An automatic driving trajectory generation method, device and terminal equipment

By acquiring autonomous driving information and using information feature fusion and Riemannian manifold constraints to generate trajectories, the redundancy and continuity issues in trajectory generation in existing technologies are solved, thereby improving the decision-making accuracy and safety of autonomous driving systems.

CN121929197BActive Publication Date: 2026-07-03JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-03-27
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as redundant text input leading to divergent focus, discretized trajectory sampling causing a lack of spatiotemporal continuity, and geometric distortion in long-term path planning. Furthermore, they suffer from significant interference in model optimization, separation of semantic reasoning and continuous control, incomplete instruction execution, lack of physical constraints in trajectory generation, and high system inference latency.

Method used

By acquiring two-dimensional image information, driving navigation command information, and three-dimensional environmental spatial information for autonomous driving, a fused feature information is generated using a preset information feature fusion model. Furthermore, a vehicle motion atomic token alignment model is used to tightly bind high-level driving intentions with low-level motion semantics. A trajectory is generated using Riemannian manifold constraints, eliminating multimodal input redundancy and achieving trajectory smoothness and safety.

Benefits of technology

It significantly improves the decision-making accuracy, trajectory smoothness, and driving safety of autonomous driving systems in complex and dynamic scenarios, while reducing inference latency and enhancing the stability of instruction execution and long-term planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121929197B_ABST
    Figure CN121929197B_ABST
Patent Text Reader

Abstract

This application provides an autonomous driving trajectory generation method, apparatus, and terminal device, applicable to the field of data processing technology. The method includes: generating autonomous driving fusion feature information based on autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and an autonomous driving information feature fusion model; obtaining aligned autonomous driving fusion feature information and aligned motion atomic token information based on the autonomous driving fusion feature information, vehicle motion atomic token information, and an alignment model of autonomous driving fusion features and motion atomic tokens; and generating autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, aligned motion atomic token information, and a target driving trajectory generation model. This application significantly improves the decision-making accuracy of autonomous driving systems in complex dynamic scenarios, while reducing inference latency and enhancing the stability of command execution and long-term planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to methods, devices and terminal equipment for generating autonomous driving trajectories. Background Technology

[0002] As end-to-end autonomous driving has become a mainstream research direction in the industry, by constructing a direct mapping from the original sensor input to the planned trajectory or control signal, the traditional perception, planning and control processes are integrated into a unified neural framework, which effectively reduces the cascading failures of modular architecture.

[0003] Current mainstream VLA driving technologies first use onboard sensors to collect environmental data and extract multimodal features. The visual features and driving command text are then jointly input into the backbone network of a visual-language model or a large language model for encoding. Discrete token generation and continuous control regression are then completed simultaneously through a shared feedforward network. Some solutions discretize vehicle actions into autoregressive tokens and combine them with flow matching or diffusion style latent variables to synthesize continuous control commands. Other technologies use sparsity strategies to optimize the computational efficiency of the visual-language model and reduce the computational load of model inference.

[0004] Existing technologies suffer from technical problems such as redundant text input leading to divergent focus, discretized trajectory sampling causing a lack of spatiotemporal continuity, and geometric distortion in long-term path planning. They also suffer from significant interference in model optimization, separation of semantic reasoning and continuous control, incomplete instruction execution, lack of physical constraints in trajectory generation, and high system inference latency. Summary of the Invention

[0005] In view of this, embodiments of this application provide an autonomous driving trajectory generation method, apparatus, and terminal device, aiming to solve the problems in the prior art such as text input redundancy causing divergence of focus, trajectory discretization sampling leading to loss of spatiotemporal continuity, and geometric distortion in long-term path planning.

[0006] The first aspect of this application provides an autonomous driving trajectory generation method, including:

[0007] Acquire two-dimensional image information, driving navigation command information, and three-dimensional environmental spatial information for autonomous driving;

[0008] Based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion model, autonomous driving fusion feature information is generated;

[0009] Based on the autonomous driving fusion feature information, the preset vehicle motion atomic token information, and the preset autonomous driving fusion feature and motion atomic token alignment model, the aligned autonomous driving fusion feature information and the aligned motion atomic token information are obtained.

[0010] Based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model, autonomous driving trajectory information is generated.

[0011] A second aspect of this application provides an autonomous driving trajectory generation apparatus, comprising:

[0012] The information acquisition module is used to acquire autonomous driving two-dimensional image information, driving navigation command information, and autonomous driving three-dimensional environmental spatial information.

[0013] The autonomous driving fusion feature information generation module is used to generate autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and a preset autonomous driving information feature fusion model.

[0014] The module for generating aligned autonomous driving fusion feature information and aligned motion atom token information is used to obtain aligned autonomous driving fusion feature information and aligned motion atom token information based on the autonomous driving fusion feature information, the preset vehicle motion atom token information and the preset alignment model of autonomous driving fusion feature and motion atom token.

[0015] The autonomous driving trajectory information generation module is used to generate autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model.

[0016] A third aspect of this application provides a terminal device, the terminal device including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the computer program to implement the steps of the autonomous driving trajectory generation method described in the first aspect above.

[0017] A fourth aspect of this application provides a computer-readable storage medium, comprising: storing a computer program, wherein when executed by a processor, the computer program implements the steps of the autonomous driving trajectory generation method described in the first aspect above.

[0018] Compared with the prior art, the beneficial effects of the embodiments of this application are: this application effectively eliminates redundant information of multimodal input, realizes the tight binding between high-level driving intention and low-level motion semantics, significantly improves the decision accuracy, trajectory smoothness and driving safety of autonomous driving system in complex dynamic scenarios, and at the same time reduces inference latency and enhances the stability of instruction execution and long-term planning. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram illustrating the implementation process of the autonomous driving trajectory generation method provided in Embodiment 1 of this application;

[0021] Figure 2 This is a schematic diagram of the implementation process of obtaining the target driving trajectory generation model provided in Embodiment 2 of this application;

[0022] Figure 3 This is a schematic diagram illustrating the implementation process of the autonomous driving trajectory generation method provided in Embodiment 3 of this application;

[0023] Figure 4 This is a schematic diagram illustrating the implementation process of the autonomous driving trajectory generation method provided in Embodiment 4 of this application;

[0024] Figure 5 This is a schematic diagram of the implementation process of the autonomous driving trajectory generation method provided in Embodiment 5 of this application;

[0025] Figure 6 This is a schematic diagram illustrating the implementation process of the autonomous driving trajectory generation method provided in Embodiment Six of this application;

[0026] Figure 7 This is a schematic diagram of the implementation process of the autonomous driving trajectory generation method provided in Embodiment 7 of this application;

[0027] Figure 8 This is a schematic diagram of the structure of the autonomous driving trajectory generation device provided in the embodiments of this application;

[0028] Figure 9 This is a schematic diagram of the terminal device provided in the embodiments of this application. Detailed Implementation

[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0030] To illustrate the technical solution described in this application, specific embodiments are provided below.

[0031] Figure 1 A flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment 1 of this application is shown, and is described in detail below:

[0032] Step S101: Obtain autonomous driving two-dimensional image information, driving navigation command information, and autonomous driving three-dimensional environmental space information.

[0033] In this embodiment, the two-dimensional image information for autonomous driving can be the visual texture, color, and target contour information of the road scene contained in the high-resolution image sequence captured by the vehicle-mounted multi-view camera, which is obtained by synchronously acquiring scene images through the vehicle-mounted multi-view camera; the driving navigation instruction information can be vehicle driving guidance information presented in natural language form, which is obtained by inputting into the system through the human-machine interface; the three-dimensional environmental spatial information for autonomous driving can be the set of three-dimensional spatial discrete points containing the positions and geometric shapes of road obstacles and traffic participants collected by the lidar, which is obtained by acquiring echo signals emitted and received by the vehicle-mounted lidar.

[0034] Step S102: Generate autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation instruction information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion model.

[0035] In this embodiment, the preset autonomous driving information feature fusion model can be pre-defined manually or constructed by combining the human visual attention system (VAS) and the LLM model. It can involve first extracting features from the autonomous driving 2D image information and the autonomous driving 3D environmental spatial information separately, then performing instruction-driven routing and sparsification processing on the extracted features based on driving navigation command information, and finally inputting the processed features into the preset autonomous driving information feature fusion model to complete deep interaction and alignment of multimodal features. Then, semantic pruning and structured compression are performed on the aligned features to improve the feature signal-to-noise ratio, thereby generating autonomous driving fused feature information.

[0036] In this embodiment, a high-resolution image sequence captured by an onboard multi-view camera is received. and driving navigation instructions in natural language (e.g., "turn left at the intersection ahead"). To eliminate redundancy in multimodal data and focus on key information, this application designs a two-stage semantic compression channel. First, it enters the visual feature aggregation stage (Stage-1), which... Input the visual encoder and introduce the visual modulation module, using instructions. Conditional modulation is applied to the visual encoding process. This process simulates a visual attention system (VAS), aggregating a large number of original patch-tokens into a small number of instruction-aware features. This reduces the number of visual tokens to 25% of the original. Then, the semantic pruning stage (Stage-2) is initiated. With instructions The data is fed into the language modulation module in parallel, where the intent relevance score obtained from language inference is used for motion intent filtering. This module removes background interference irrelevant to the current driving intent (such as roadside billboards) and outputs extremely sparse and semantically highly condensed key features. This provides a high signal-to-noise ratio input for subsequent processing. It is used to obtain sparse semantic features. Subsequently, to achieve effective connection between high-level intent and low-level control, a feature alignment and discretization representation module was designed. Specifically, firstly, the cross-modal fusion module receives... By employing a multi-head attention mechanism to deeply interact with visual and linguistic information, aligned and fused features are generated. Next, a pre-built discretized kinematics codebook is introduced. Each element represents a basic motion primitive, containing specific curvature, acceleration, and steering angle states. The system continuously fuses these features. Mapped to this discrete space, sequence motion atomic tokens are generated through quantization operations. This process transforms abstract driving intentions into concrete sequences of actions. While preserving the fine-grained kinematic information required for continuous control, it achieves structured compression of the feature space.

[0037] Specifically, firstly, a visual encoder (such as ViT) is used to divide image I into N patches to obtain an initial visual token sequence. Mapped to instruction vector To simulate the human visual attention system (VAS), an instruction-gated sparse attention mechanism (IGSA) can be used to calculate the attention score of the instruction on the visual token. :

[0038]

[0039]

[0040] Based on attention scores, hard mask sparsification is performed, retaining the Top-K key tokens to generate sparse visual features.

[0041]

[0042] in, For dynamic thresholds, This is an indicator function. Subsequently, and The input semantic pruning module is further fused through a cross-attention mechanism to output a high signal-to-noise ratio semantic feature vector. ,in This represents the length of the compressed sequence. This process compresses the number of visual tokens to less than 25% of the original, which can be expressed as:

[0043]

[0044] The two-stage compression avoids losing critical information because Stage-1 primarily performs "continuous suppression to reduce redundancy" and preserves the anchor token to ensure that critical security elements always exist; Stage-2 performs "hard selection + global digest retrieval," and through... A dynamic threshold mechanism ensures that a sufficient number of interactive entities and topological semantics are retained in complex scenarios, thereby significantly reducing token computation while maintaining the key semantic and geometric constraint information required for driving decisions. The preferred anchor token set includes: vehicle state token, lane / drivable area token, nearest neighbor traffic participant token, and traffic rule token; for pruned tokens, a weighted aggregation method is preferred to obtain a global summary token. It is then concatenated with the retained token and input for subsequent cross-modal fusion, thereby achieving information retrieval; at the same time, a minimum retention number is set. (or make the Top-K ratio) (Not lower than the preset lower limit) to avoid excessive compression in complex interactive scenarios.

[0045] Step S103: Based on the autonomous driving fusion feature information, the preset vehicle motion atomic token information, and the preset autonomous driving fusion feature and motion atomic token alignment model, the aligned autonomous driving fusion feature information and the aligned motion atomic token information are obtained.

[0046] In this embodiment, the preset vehicle motion atomic token information can be manually preset, and the preset alignment model between the autonomous driving fusion features and the motion atomic tokens can also be manually preset, and can be constructed based on a multi-head attention mechanism and vector quantization strategy. The process can begin by inputting the autonomous driving fusion feature information into the preset alignment model between the autonomous driving fusion features and the motion atomic tokens, then aligning the feature space of the autonomous driving fusion feature information and the preset vehicle motion atomic token information through cross-modal attention interaction. Next, the aligned features are discretized and fine-grained kinematic information is preserved. Finally, stability constraints and sequence smoothing regularization are applied to the mapped information to obtain the aligned autonomous driving fusion feature information and the aligned motion atomic token information.

[0047] In this embodiment, the preset vehicle motion atomic token information can be obtained by first collecting the driving trajectories of multiple autonomous vehicles, and then classifying the coordinate points of the driving trajectories of multiple autonomous vehicles to obtain 2048 types of coordinate points as vehicle motion atomic token information. That is, all types of autonomous driving trajectories can be represented by these 2048 types of coordinate points.

[0048] In this embodiment, based on the generated motion atom tokens Trajectory generation is performed within an implicit neural manifold field. A Riemannian manifold space embedded with vehicle non-holonomic constraints is constructed. Within this space, the trajectory generation process is modeled as a manifold evolution constrained by Riemannian geometry, in order to... The state transition is extrapolated, and the manifold trajectory features containing spatiotemporal information evolve along the geodesic direction. This allows the topological properties of the trajectory to be explicitly constrained using the geometry of the manifold, ensuring the generated features are... It naturally satisfies physical limitations such as the minimum turning radius of vehicles, thereby ensuring the geometric continuity and consistency of long-term planning.

[0049] Specifically, multi-head attention mechanisms can be used to integrate semantic features. Mapped to fusion features :

[0050]

[0051] Pre-construct a discretized kinematic codebook , where each code vector Represents a basic kinematic primitive (containing a specific curvature) Vector quantization is used to fuse continuous features. Mapping to discrete space. Features for time step t. Find the nearest codebook vector :

[0052] The generated motion atom token sequence is .

[0053] To ensure the stability of the quantization process, commitment loss is introduced. To prevent encoder divergence:

[0054]

[0055] Here, sg[⋅] represents the stop-gradient operation, and β is the weighting coefficient. This step ensures that the abstract driving intention is transformed into a concrete, physically interpretable combination of action sequences.

[0056] Motion atomic tokens transform the planning conditions from a noisy, continuous representation to a smooth evolution in a finite state space through "VQ discrete stabilization + primitive reuse + temporal regularization," thereby improving the temporal consistency and controllability of the output trajectory and control quantity, and reducing short-term jitter and cross-frame discontinuities. Preferably, an atomic token sequence smoothing regularization term is added to the training objective. This is used to penalize frequent switching of atoms in adjacent time steps (or to penalize differences in atom embedding), and its weight is... This is to further reduce trajectory jitter and cross-frame discontinuity.

[0057] Step S104: Generate autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, aligned motion atom token information, and target driving trajectory generation model.

[0058] In this embodiment, the target driving trajectory generation model can be pre-set by humans or constructed by embedding vehicle non-integrity constraints and Riemannian geometric rules. It can be implemented by first inputting aligned autonomous driving fusion feature information and aligned motion atom token information into the target driving trajectory generation model, then performing trajectory state deduction and geodesic direction evolution in an implicit neural manifold field, subsequently applying geometric topological and kinematic feasibility constraints to the deduced trajectory features, and finally mapping the constrained trajectory features back to Euclidean space to complete path calculation, thereby generating autonomous driving trajectory information.

[0059] In this embodiment, the vehicle state space is defined as a Riemannian manifold. The Riemannian metric is defined on it. This manifold incorporates non-holonomic constraints for the vehicle. For the state... Its tangent space is Constructing implicit neural manifold fields Trajectory generation is modeled as a manifold evolution process constrained by Riemannian geometry.

[0060] Based on motion atomic tokens Given that the trajectory evolution follows the geodesic equation, let the trajectory be γ(t), and its evolution satisfy the second-order differential equation:

[0061]

[0062] in, Decide:

[0063]

[0064] The target driving trajectory generation model uses a neural network to parameterize the metric tensor. This ensures that the generated trajectory γ(t) naturally satisfies physical constraints such as the vehicle's minimum turning radius.

[0065] To train the manifold evolution field, a flow matching loss on the manifold is defined. Let the target trajectory distribution be... The prior distribution is The vector field is :

[0066]

[0067] in, This represents the norm based on the Riemannian metric. This method explicitly constrains the topological properties of the trajectory using the geometry of the manifold.

[0068] Riemannian manifolds incorporate physical feasibility into the trajectory generation process through "learnable local geometric cost (metric) + continuity constraints of manifold evolution + nonholonomic consistency," thereby fundamentally suppressing geometric distortions such as sharp turns and curvature jumps, and improving safety and comfort.

[0069] The autonomous driving trajectory generation method provided in this application effectively eliminates redundant information from multimodal inputs, achieves a tight binding between high-level driving intentions and low-level motion semantics, significantly improves the decision-making accuracy, trajectory smoothness, and driving safety of the autonomous driving system in complex dynamic scenarios, while reducing inference latency and enhancing the stability of instruction execution and long-term planning.

[0070] Figure 2 The flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment 2 of this application is shown. The difference between this method and Embodiment 1 is that the target driving trajectory generation model is obtained through the following steps:

[0071] Step S201: Obtain autonomous driving visual perception information, autonomous driving semantic command information, autonomous driving spatial environment perception information, historical driving trajectory information, and initial driving trajectory generation model.

[0072] In this embodiment, the autonomous driving visual perception information can be road scene image feature information collected by an onboard multi-view camera and extracted by an onboard visual encoder; the autonomous driving semantic instruction information can be driving guidance information in natural language form and encoded by a language encoder; the autonomous driving spatial environment perception information can be three-dimensional spatial point cloud feature information collected by an onboard LiDAR and extracted by a LiDAR feature encoder; the historical driving trajectory information can be the path points and control state sequence information of the vehicle's past travels, which can be used to represent the vehicle's actual driving trajectory and can be collected by an onboard driving recording unit; the initial driving trajectory generation model can be preset by humans or an untrained trajectory generation network model built based on neural manifold structures and attention mechanisms. It is understood that the autonomous driving visual perception information, autonomous driving semantic instruction information, autonomous driving spatial environment perception information, and historical driving trajectory information are all data used to train the initial driving trajectory generation model.

[0073] Step S202: Based on the autonomous driving visual perception information, autonomous driving semantic instruction information, autonomous driving spatial environment perception information, historical driving trajectory information, and preset vehicle motion atomic token information, the initial driving trajectory generation model is trained to obtain the target driving trajectory generation model.

[0074] In this embodiment, the preset vehicle motion atomic token information can be manually preset or can be discretized kinematic representation information containing multiple basic motion primitives. Autonomous driving visual perception information, autonomous driving semantic command information, autonomous driving spatial environment perception information, and the preset vehicle motion atomic token information can be input into the initial driving trajectory generation model to complete forward inference. Then, using historical driving trajectory information as a supervision signal, the difference between the model output and the real trajectory is calculated. Finally, based on the difference value, the network parameters of the initial driving trajectory generation model are iteratively updated, thereby completing model training and generating the target driving trajectory generation model.

[0075] The autonomous driving trajectory generation method provided in this application improves the adaptability of the target driving trajectory generation model to driving scenarios and the accuracy of trajectory output by jointly training the model with multimodal perception information and historical driving trajectory information, so that the output of the target driving trajectory generation model is more in line with the physical constraints and command requirements of real driving.

[0076] Figure 3The flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment 3 of this application is shown. The difference between this method and Embodiment 2 is that step S202 specifically includes:

[0077] Step S301: Generate autonomous driving trajectory information to be confirmed based on the autonomous driving visual perception information, autonomous driving semantic instruction information, autonomous driving spatial environment perception information, preset vehicle motion atomic token information, and initial driving trajectory generation model.

[0078] In this embodiment, the preset vehicle motion atomic token information can be manually preset. It can be achieved by sequentially inputting autonomous driving visual perception information, autonomous driving semantic command information, autonomous driving spatial environment perception information, and the preset vehicle motion atomic token information into an initial driving trajectory generation model. Then, trajectory deduction is completed through cross-modal fusion and manifold evolution calculations within the model, and finally, an unverified and uncorrected trajectory sequence is output, thereby generating autonomous driving trajectory information to be confirmed.

[0079] Step S302: Calculate the basic loss information of the driving trajectory based on the autonomous driving trajectory information to be confirmed and the historical driving trajectory information.

[0080] In this embodiment, the autonomous driving trajectory information to be confirmed is compared with the historical driving trajectory information on a time-by-time basis. Then, the deviation values ​​of the two types of trajectories in terms of position and state are calculated. The deviation values ​​are then normalized and weighted to generate basic loss information of the driving trajectory.

[0081] Step S303: Based on the autonomous driving trajectory information to be confirmed and multiple preset driving trajectory loss calculation functions, calculate the driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, and driving trajectory task completion loss information.

[0082] In this embodiment, the multiple preset driving trajectory loss calculation functions can be manually preset. The autonomous driving trajectory information to be confirmed can be input into these preset functions, and then the trajectory geometric compliance deviation, trajectory curvature and acceleration change deviation, and instruction execution compliance deviation can be calculated using different functions. The three independent deviation results are then extracted to generate driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, and driving trajectory task completion loss information.

[0083] Step S304: Based on the driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, driving trajectory task completion loss information, driving trajectory basic loss information, and multiple preset autonomous driving trajectory loss calculation weights, calculate the autonomous driving trajectory composite loss information.

[0084] In this embodiment, the multiple preset weights for calculating autonomous driving trajectory loss can be manually preset. The four types of loss information can be multiplied by the corresponding multiple preset weights for calculating autonomous driving trajectory loss, and then the weighted loss values ​​are summed and summarized. Finally, the overall loss is normalized to generate composite loss information for autonomous driving trajectory.

[0085] Step S305: Based on the initial driving trajectory generation model, the autonomous driving trajectory composite loss information, and the preset autonomous driving trajectory composite loss threshold, the target driving trajectory generation model is obtained.

[0086] In this embodiment, the preset composite loss threshold for autonomous driving trajectory can be manually preset. It can be achieved by comparing the composite loss information of the autonomous driving trajectory with the preset composite loss threshold, then iteratively optimizing the parameters of the initial driving trajectory generation model through backpropagation based on the comparison results, and then repeating the training until the loss value meets the requirements of the preset composite loss threshold for autonomous driving trajectory, thereby generating the target driving trajectory generation model.

[0087] The autonomous driving trajectory generation method provided in this application enhances the geometric rationality, motion smoothness, and command execution accuracy of trajectory generation by training an autonomous driving trajectory composite loss information constraint model, thereby improving the reliability and safety of autonomous driving trajectory output.

[0088] Figure 4 The following is a flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment 4 of this application. The difference between this method and Embodiment 3 described above is that:

[0089] Multiple preset driving trajectory loss calculation functions include preset driving trajectory geometric constraint loss calculation function, preset driving trajectory smoothness loss calculation function, and preset driving trajectory task completion loss calculation function;

[0090] Step S303 specifically includes:

[0091] Step S401: Calculate the driving trajectory geometric constraint loss information based on the autonomous driving trajectory information to be confirmed and the preset driving trajectory geometric constraint loss calculation function.

[0092] In this embodiment, the preset driving trajectory geometric constraint loss calculation function can be manually preset. The autonomous driving trajectory information to be confirmed can be input into the preset driving trajectory geometric constraint loss calculation function, which then detects whether the trajectory deviates from the feasible manifold surface and the vehicle's physical constraint range. The degree of trajectory violation of geometric constraints is then statistically analyzed to generate driving trajectory geometric constraint loss information, which is used to penalize the state of deviating from the manifold surface and ensure kinematic feasibility.

[0093] In this embodiment, the preset driving trajectory geometric constraint loss calculation function can be expressed as: :

[0094]

[0095] Step S402: Calculate the driving trajectory smoothness loss information based on the autonomous driving trajectory information to be confirmed and the preset driving trajectory smoothness loss calculation function.

[0096] In this embodiment, the preset driving trajectory smoothness loss calculation function can be manually preset. The autonomous driving trajectory information to be confirmed can be input into the preset driving trajectory smoothness loss calculation function, which then calculates the curvature and acceleration change amplitude of adjacent time points of the trajectory, and then statistically analyzes the degree of trajectory fluctuation and jump, thereby generating driving trajectory smoothness loss information to constrain the rate of curvature change and ensure ride comfort.

[0097] In this embodiment, the preset driving trajectory smoothness loss calculation function can be expressed as: :

[0098]

[0099] Step S403: Calculate the driving trajectory task completion loss information based on the autonomous driving trajectory information to be confirmed and the preset driving trajectory task completion loss calculation function.

[0100] In this embodiment, the preset driving trajectory task completion loss calculation function can be manually preset. The autonomous driving trajectory information to be confirmed can be input into the preset driving trajectory task completion loss calculation function, and then the deviation between the trajectory endpoint and the target position is compared and the instruction execution status is statistically analyzed. Then, the loss value caused by the task not being completed is calculated, thereby generating driving trajectory task completion loss information, which is used to verify the instruction execution status (such as endpoint error).

[0101] In this embodiment, the preset driving trajectory task completion loss calculation function can be expressed as: :

[0102]

[0103] In this embodiment, the formula for calculating the composite loss information of the autonomous driving trajectory can be expressed as:

[0104]

[0105] Multiple preset weights for calculating autonomous driving trajectory loss can be set to values ​​of The training scheme uniformly adopts AdamW, with end-to-end training completed in 20 epochs. The learning rate is set to 9e-5, and a minimum learning rate constraint of 5e-5 is used to ensure stable updates in later stages.

[0106] The autonomous driving trajectory generation method provided in this application accurately locates the source of deviation in the training of the target driving trajectory generation model, making the loss constraint more targeted and interpretable, thereby improving the quality of autonomous driving trajectory generation.

[0107] Figure 5 The flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment 5 of this application is shown. Its difference from Embodiment 1 described above lies in:

[0108] The preset autonomous driving information feature fusion model includes a preset autonomous driving visual feature aggregation sub-model, a preset autonomous driving semantic feature generation sub-model, and a preset autonomous driving information feature fusion sub-model.

[0109] Step S102 specifically includes:

[0110] Step S501: Generate autonomous driving visual feature aggregation information based on the autonomous driving two-dimensional image information and the preset autonomous driving visual feature aggregation sub-model.

[0111] In this embodiment, the preset autonomous driving visual feature aggregation sub-model can be manually preset or constructed by combining a visual encoder and a command-gated sparse attention mechanism. It can involve inputting autonomous driving two-dimensional image information into the preset autonomous driving visual feature aggregation sub-model, then performing block encoding and command condition modulation on the image information, followed by continuous suppression and redundancy reduction processing while retaining safety-critical anchor point features, and finally performing global summary retrieval and dynamic threshold filtering on the features to generate autonomous driving visual feature aggregation information.

[0112] Step S502: Generate autonomous driving semantic feature information based on the autonomous driving visual feature aggregation information, driving navigation instruction information, and the preset autonomous driving semantic feature generation sub-model.

[0113] In this embodiment, the preset autonomous driving semantic feature generation sub-model can be manually preset or constructed by combining a language modulation module and a cross-attention mechanism. It can input autonomous driving visual feature aggregation information and driving navigation command information into the preset autonomous driving semantic feature generation sub-model, then calculate the intent relevance score and complete motion intent filtering, subsequently removing background interference features irrelevant to the driving task, and then highly condensing and sparsifying the features to generate autonomous driving semantic feature information.

[0114] Step S503: Generate autonomous driving fusion feature information based on the autonomous driving semantic feature information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion sub-model.

[0115] In this embodiment, the preset autonomous driving information feature fusion sub-model can be manually preset, or it can be a cross-modal fusion network built based on a multi-head attention mechanism. It can input autonomous driving semantic feature information and autonomous driving 3D environmental spatial information into the preset autonomous driving information feature fusion sub-model to achieve deep interactive alignment of visual language features and spatial environment features. Then, the aligned features are subjected to structured compression and fine-grained information preservation, and finally, high signal-to-noise ratio multimodal integrated features are output, thereby generating autonomous driving fused feature information.

[0116] The autonomous driving trajectory generation method provided in this application is used to achieve visual aggregation, semantic pruning and multimodal fusion, thereby improving the accuracy and efficiency of feature extraction, effectively reducing input redundancy and enhancing the retention of key driving information.

[0117] Figure 6 The following is a flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment Six of this application. The difference between this method and Embodiment One described above is that:

[0118] The preset autonomous driving fusion feature and motion atom token alignment model includes a preset autonomous driving visual language feature generation sub-model and a preset autonomous driving action feature generation sub-model.

[0119] Step S103 specifically includes:

[0120] Step S601: Obtain autonomous driving noise information.

[0121] In this embodiment, the autonomous driving noise information can be used to train the target driving trajectory generation model to improve its denoising capability. It is obtained through a random sampling module sampling according to a preset distribution, which can be randomly generated according to a Gaussian distribution. Denoising refers to continuously removing noise multiple times from a Gaussian-distributed noise source to obtain continuous, accurate rules or strategies. Motion atomic tokens can help the target driving trajectory generation model understand the autonomous driving task during training, thereby improving its denoising capability.

[0122] Step S602: Based on the autonomous driving fusion feature information and the preset autonomous driving visual language feature generation sub-model, the aligned autonomous driving fusion feature information is obtained.

[0123] In this embodiment, the preset autonomous driving visual language feature generation sub-model can be manually preset or constructed based on a cross-modal attention mechanism. It can involve inputting autonomous driving fused feature information into the preset autonomous driving visual language feature generation sub-model, then refining the features through feature space mapping and semantic alignment, followed by stability constraints and information enhancement processing to obtain aligned autonomous driving fused feature information.

[0124] Step S603: Generate a sub-model based on the autonomous driving noise information and the preset autonomous driving action features to obtain aligned motion atom token information.

[0125] In this embodiment, the preset autonomous driving action feature generation sub-model can be manually preset, or it can be constructed based on vector quantization and kinematic code. It can involve inputting autonomous driving noise information into the preset autonomous driving action feature generation sub-model, then combining discretized kinematic priors to initialize and map motion atomic tokens, and subsequently performing smoothing regularization and stability constraint processing on the token sequence to obtain aligned motion atomic token information.

[0126] The autonomous driving trajectory generation method provided in this application improves the accuracy of multimodal feature alignment and the stability of motion atomic token generation, thereby ensuring the continuity and feasibility of subsequent autonomous driving trajectory generation.

[0127] Figure 7 The following is a flowchart illustrating the implementation of the autonomous driving trajectory generation method provided in Embodiment Seven of this application. The difference between this method and Embodiment One described above is that:

[0128] The autonomous driving trajectory information includes autonomous driving trajectory environmental perception information, autonomous driving trajectory strategy information, driving trajectory vehicle status information, and autonomous driving trajectory prediction information;

[0129] Step S104 specifically includes:

[0130] Step S701: Based on the aligned autonomous driving fusion feature information and the target driving trajectory generation model, generate autonomous driving trajectory environment perception information, autonomous driving trajectory strategy information, and driving trajectory vehicle state information.

[0131] In this embodiment, the target driving trajectory generation model can be pre-set by humans or constructed by embedding vehicle non-integrity constraints and Riemannian geometric rules. Aligned autonomous driving fusion feature information can be input into the target driving trajectory generation model, which then performs multimodal semantic feature parsing and state deduction within an implicit neural manifold field. Subsequently, environmental topological features, driving decision features, and vehicle state features are extracted based on manifold geometric constraints. These three types of features are then independently decoded and normalized for output, thereby generating autonomous driving trajectory environmental perception information, autonomous driving trajectory strategy information, and vehicle state information.

[0132] Step S702: Generate autonomous driving trajectory prediction information based on the aligned motion atom token information and the target driving trajectory generation model.

[0133] In this embodiment, the target driving trajectory generation model can be pre-set by the user or constructed by embedding vehicle non-integrity constraints and Riemannian geometric rules. Aligned motion atom token information can be input into the target driving trajectory generation model, and then the trajectory can be evolved along the geodesic direction within the Riemannian manifold space under the condition of the motion atom tokens. The evolution results are then subject to kinematic feasibility constraints and geometric topological regularization. Finally, the manifold space features are mapped back to Euclidean space to complete the path point sequence calculation, thereby generating autonomous driving trajectory prediction information.

[0134] The autonomous driving trajectory generation method provided in this application is used to achieve deep collaboration between trajectory generation and scene understanding, decision planning, and state perception, thereby improving the decision reliability and trajectory execution accuracy of the autonomous driving system in complex dynamic scenarios.

[0135] Corresponding to the method in the above embodiments, Figure 8 A structural block diagram of an autonomous driving trajectory generation device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown. Figure 8 The example autonomous driving trajectory generation device can be the execution subject of the autonomous driving trajectory generation method provided in the aforementioned embodiment 1.

[0136] Reference Figure 8 The autonomous driving trajectory generation device includes:

[0137] The information acquisition module 810 is used to acquire autonomous driving two-dimensional image information, driving navigation command information, and autonomous driving three-dimensional environmental spatial information.

[0138] The autonomous driving fusion feature information generation module 820 is used to generate autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental space information, and a preset autonomous driving information feature fusion model.

[0139] The aligned autonomous driving fusion feature information and aligned motion atom token information generation module 830 is used to obtain aligned autonomous driving fusion feature information and aligned motion atom token information based on the autonomous driving fusion feature information, the preset vehicle motion atom token information and the preset autonomous driving fusion feature and motion atom token alignment model.

[0140] The autonomous driving trajectory information generation module 840 is used to generate autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model.

[0141] For details on how each module in the autonomous driving trajectory generation device provided in this application implements its respective function, please refer to the foregoing. Figure 1 The description of Embodiment 1 shown will not be repeated here.

[0142] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0143] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0144] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0145] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0146] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first," "second," etc., are used in the text to describe various elements in some embodiments of this application, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first table may be named a second table, and similarly, a second table may be named a first table, without departing from the scope of the various described embodiments. Both the first table and the second table are tables, but they are not the same table.

[0147] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0148] The autonomous driving trajectory generation method provided in this application can be applied to terminal devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality / virtual reality devices, laptops, super mobile personal computers, netbooks, and personal digital assistants. This application does not impose any restrictions on the specific type of terminal device.

[0149] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. For example... Figure 9 As shown, the terminal device 9 of this embodiment includes: at least one processor 90 ( Figure 9 (Only one is shown in the image) a memory 91, which stores a computer program 92 that can run on the processor 90. When the processor 90 executes the computer program 92, it implements the steps in the various embodiments of the autonomous driving trajectory generation method described above, for example... Figure 1 Steps S101 to S104 are shown. Alternatively, when the processor 90 executes the computer program 92, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 8 The functions of the information acquisition module 810, the autonomous driving fusion feature information generation module 820, the aligned autonomous driving fusion feature information and aligned motion atom token information generation module 830, and the autonomous driving trajectory information generation module 840 are shown.

[0150] The terminal device 9 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor 90 and a memory 91. Those skilled in the art will understand that... Figure 9 This is merely an example of terminal device 9 and does not constitute a limitation on terminal device 9. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input transmission devices, network access devices, buses, etc.

[0151] The processor 90 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0152] In some embodiments, the memory 91 may be an internal storage unit of the terminal device 9, such as a hard disk or memory of the terminal device 9. The memory 91 may also be an external storage device of the terminal device 9, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the terminal device 9. Furthermore, the memory 91 may include both internal and external storage units of the terminal device 9. The memory 91 is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of computer programs. The memory 91 can also be used to temporarily store data that has been sent or will be sent.

[0153] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0154] This application also provides a terminal device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, it causes the terminal device to implement the steps in any of the above method embodiments.

[0155] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0156] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0157] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0158] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0159] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating autonomous driving trajectories, characterized in that, include: Acquire two-dimensional image information, driving navigation command information, and three-dimensional environmental spatial information for autonomous driving; Based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion model, autonomous driving fusion feature information is generated; Based on the autonomous driving fusion feature information, the preset vehicle motion atomic token information, and the preset autonomous driving fusion feature and motion atomic token alignment model, the aligned autonomous driving fusion feature information and the aligned motion atomic token information are obtained. Based on the aligned autonomous driving fusion feature information, aligned motion atom token information, and target driving trajectory generation model, generate autonomous driving trajectory information; The preset autonomous driving information feature fusion model includes a preset autonomous driving visual feature aggregation sub-model, a preset autonomous driving semantic feature generation sub-model, and a preset autonomous driving information feature fusion sub-model. The step of generating autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and a preset autonomous driving information feature fusion model specifically includes: Based on the autonomous driving two-dimensional image information and the preset autonomous driving visual feature aggregation sub-model, autonomous driving visual feature aggregation information is generated; Based on the autonomous driving visual feature aggregation information, driving navigation instruction information, and the preset autonomous driving semantic feature generation sub-model, autonomous driving semantic feature information is generated; Based on the autonomous driving semantic feature information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion sub-model, autonomous driving fusion feature information is generated; The preset autonomous driving fusion feature and motion atom token alignment model includes a preset autonomous driving visual language feature generation sub-model and a preset autonomous driving action feature generation sub-model. The step of obtaining aligned autonomous driving fusion feature information and aligned motion atomic token information based on the autonomous driving fusion feature information, preset vehicle motion atomic token information, and preset autonomous driving fusion feature and motion atomic token alignment model specifically includes: Obtain noise information for autonomous driving; Based on the autonomous driving fusion feature information and the preset autonomous driving visual language feature generation sub-model, the aligned autonomous driving fusion feature information is obtained. Based on the autonomous driving noise information and the preset autonomous driving action features, a sub-model is generated to obtain aligned motion atom token information; The autonomous driving trajectory information includes autonomous driving trajectory environmental perception information, autonomous driving trajectory strategy information, driving trajectory vehicle status information, and autonomous driving trajectory prediction information; The step of generating autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model specifically includes: Based on the aligned autonomous driving fusion feature information and the target driving trajectory generation model, autonomous driving trajectory environment perception information, autonomous driving trajectory strategy information and driving trajectory vehicle state information are generated. Based on the aligned motion atomic token information and the target driving trajectory generation model, autonomous driving trajectory prediction information is generated.

2. The autonomous driving trajectory generation method as described in claim 1, characterized in that, The target driving trajectory generation model is obtained through the following steps: Acquire autonomous driving visual perception information, autonomous driving semantic command information, autonomous driving spatial environment perception information, historical driving trajectory information, and initial driving trajectory generation model; Based on the autonomous driving visual perception information, autonomous driving semantic instruction information, autonomous driving spatial environment perception information, historical driving trajectory information, and preset vehicle motion atomic token information, the initial driving trajectory generation model is trained to obtain the target driving trajectory generation model.

3. The autonomous driving trajectory generation method as described in claim 2, characterized in that, The step of training the initial driving trajectory generation model based on the autonomous driving visual perception information, autonomous driving semantic command information, autonomous driving spatial environment perception information, historical driving trajectory information, and preset vehicle motion atomic token information to obtain the target driving trajectory generation model specifically includes: Based on the autonomous driving visual perception information, autonomous driving semantic instruction information, autonomous driving spatial environment perception information, preset vehicle motion atomic token information, and initial driving trajectory generation model, generate autonomous driving trajectory information to be confirmed. Based on the unconfirmed autonomous driving trajectory information and the historical driving trajectory information, the basic loss information of the driving trajectory is calculated; Based on the unconfirmed autonomous driving trajectory information and multiple preset driving trajectory loss calculation functions, driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, and driving trajectory task completion loss information are calculated. Based on the driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, driving trajectory task completion loss information, driving trajectory basic loss information, and multiple preset autonomous driving trajectory loss calculation weights, the autonomous driving trajectory composite loss information is calculated. Based on the initial driving trajectory generation model, the autonomous driving trajectory composite loss information, and the preset autonomous driving trajectory composite loss threshold, the target driving trajectory generation model is obtained.

4. The autonomous driving trajectory generation method as described in claim 3, characterized in that, Multiple preset driving trajectory loss calculation functions include preset driving trajectory geometric constraint loss calculation function, preset driving trajectory smoothness loss calculation function, and preset driving trajectory task completion loss calculation function; The step of calculating driving trajectory geometric constraint loss information, driving trajectory smoothness loss information, and driving trajectory task completion loss information based on the unconfirmed autonomous driving trajectory information and multiple preset driving trajectory loss calculation functions specifically includes: Based on the autonomous driving trajectory information to be confirmed and the preset driving trajectory geometric constraint loss calculation function, the driving trajectory geometric constraint loss information is calculated. Based on the unconfirmed autonomous driving trajectory information and the preset driving trajectory smoothness loss calculation function, the driving trajectory smoothness loss information is calculated. Based on the unconfirmed autonomous driving trajectory information and the preset driving trajectory task completion loss calculation function, the driving trajectory task completion loss information is calculated.

5. An autonomous driving trajectory generation device, characterized in that, include: The information acquisition module is used to acquire autonomous driving two-dimensional image information, driving navigation command information, and autonomous driving three-dimensional environmental spatial information. The autonomous driving fusion feature information generation module is used to generate autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and a preset autonomous driving information feature fusion model. The module for generating aligned autonomous driving fusion feature information and aligned motion atom token information is used to obtain aligned autonomous driving fusion feature information and aligned motion atom token information based on the autonomous driving fusion feature information, the preset vehicle motion atom token information and the preset alignment model of autonomous driving fusion feature and motion atom token. The autonomous driving trajectory information generation module is used to generate autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model. The preset autonomous driving information feature fusion model includes a preset autonomous driving visual feature aggregation sub-model, a preset autonomous driving semantic feature generation sub-model, and a preset autonomous driving information feature fusion sub-model. The step of generating autonomous driving fusion feature information based on the autonomous driving two-dimensional image information, driving navigation command information, autonomous driving three-dimensional environmental spatial information, and a preset autonomous driving information feature fusion model specifically includes: Based on the autonomous driving two-dimensional image information and the preset autonomous driving visual feature aggregation sub-model, autonomous driving visual feature aggregation information is generated; Based on the autonomous driving visual feature aggregation information, driving navigation instruction information, and the preset autonomous driving semantic feature generation sub-model, autonomous driving semantic feature information is generated; Based on the autonomous driving semantic feature information, autonomous driving three-dimensional environmental spatial information, and the preset autonomous driving information feature fusion sub-model, autonomous driving fusion feature information is generated; The preset autonomous driving fusion feature and motion atom token alignment model includes a preset autonomous driving visual language feature generation sub-model and a preset autonomous driving action feature generation sub-model. The step of obtaining aligned autonomous driving fusion feature information and aligned motion atomic token information based on the autonomous driving fusion feature information, preset vehicle motion atomic token information, and preset autonomous driving fusion feature and motion atomic token alignment model specifically includes: Obtain noise information for autonomous driving; Based on the autonomous driving fusion feature information and the preset autonomous driving visual language feature generation sub-model, the aligned autonomous driving fusion feature information is obtained. Based on the autonomous driving noise information and the preset autonomous driving action features, a sub-model is generated to obtain aligned motion atom token information; The autonomous driving trajectory information includes autonomous driving trajectory environmental perception information, autonomous driving trajectory strategy information, driving trajectory vehicle status information, and autonomous driving trajectory prediction information; The step of generating autonomous driving trajectory information based on the aligned autonomous driving fusion feature information, the aligned motion atom token information, and the target driving trajectory generation model specifically includes: Based on the aligned autonomous driving fusion feature information and the target driving trajectory generation model, autonomous driving trajectory environment perception information, autonomous driving trajectory strategy information and driving trajectory vehicle state information are generated. Based on the aligned motion atomic token information and the target driving trajectory generation model, autonomous driving trajectory prediction information is generated.

6. A terminal device, characterized in that, The terminal device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Automatic driving double-reconstruction-mode motion planning method and device fusing visual language model, and medium

    CN121214408A

  • VLM model intelligent decision-making-based driving method and device, and storage medium

    CN121291416A

  • Automatic driving method, device and system, model fine tuning method, device and system and vehicle

    CN121404318A