Track planning model generation device and method, vehicle track planning device and method and medium

By employing a multi-stage training and closed-loop interactive feedback approach, the consistency problem between the visual language model and the end-to-end model in trajectory planning during intelligent driving was resolved. This improved the robustness and accuracy of the trajectory planning model, thereby enhancing the driving safety of intelligent vehicles.

CN121786477APending Publication Date: 2026-04-03HORIZON JOURNEY TAGE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In intelligent driving, the lack of a unified behavioral semantic alignment mechanism between visual language models and end-to-end models leads to low robustness and accuracy in vehicle trajectory prediction, affecting the safety of vehicle driving control.

Method used

By training the initial trajectory planning model in multiple stages and combining training samples from multiple dimensions, closed-loop interactive feedback is introduced to improve the output consistency between the visual language model and the end-to-end model, thereby enhancing the robustness and accuracy of the trajectory planning model.

Benefits of technology

It improves the robustness and accuracy of trajectory planning models, enhances the driving safety of intelligent vehicles, and ensures accurate trajectory prediction of vehicles in real traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786477A_ABST
    Figure CN121786477A_ABST
Patent Text Reader

Abstract

The invention discloses a trajectory planning model generation device and method, a vehicle trajectory planning device and method and a medium, and the method comprises the steps: obtaining a training sample set, enabling a training sample to comprise at least one frame of input data and at least one frame of label data, enabling the input data to comprise environment perception data, agent state data and interaction data, enabling the label data to comprise a trajectory label and a decision label, and based on the training sample set, performing multi-stage training on the initial trajectory planning model to obtain the trajectory planning model, thereby solving the problem that a decision generated by a visual language model in the trajectory planning model is inconsistent with a trajectory planned by an end-to-end model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to intelligent driving technology, and in particular to a trajectory planning model generation, a vehicle trajectory planning device, method, and medium. Background Technology

[0002] During the operation of an intelligent driving vehicle, based on navigation information and vehicle perception information, the vehicle's driving trajectory in the future time period (such as the next 6 seconds) can be planned to control the vehicle's safe driving.

[0003] In related technologies, a scheme combining a Vision Language Model (VLM) with an end-to-end (E2E) model is used for trajectory planning to obtain the vehicle's driving trajectory over a future time period. However, due to the lack of a unified behavioral semantic alignment mechanism between the VLM and the E2E model, mismatches easily arise in real-world traffic environments between the decisions generated by the VLM and the planned trajectories generated by the E2E model. This results in low robustness and accuracy of vehicle trajectory prediction, negatively impacting the safety of vehicle driving control. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a trajectory planning model generation, vehicle trajectory planning apparatus, method, and medium to solve the problem of low robustness and accuracy in vehicle trajectory prediction during intelligent driving.

[0005] A first aspect of this disclosure provides an apparatus for generating a trajectory planning model, including a processor configured to:

[0006] Obtain a training sample set, wherein the training sample includes at least one frame of input data and at least one frame of label data, wherein the input data includes environmental perception data, agent state data, and interaction data, and the label data includes trajectory labels and decision labels;

[0007] Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model.

[0008] A second aspect of this disclosure provides a vehicle trajectory planning apparatus, including a processor configured to:

[0009] Acquire vehicle environmental perception data, vehicle status data, and interaction data;

[0010] Based on the vehicle status data and the interaction data, determine the text command;

[0011] Based on the first sub-model in the trajectory planning model, the environmental perception data and the text instructions are processed to obtain the first multimodal feature;

[0012] Based on the first multimodal features and decision information features, the decision features are determined;

[0013] Based on the second sub-model in the trajectory planning model, the decision features, the environmental perception data, and the vehicle state data are processed to obtain the vehicle's driving trajectory in the future time period.

[0014] A third aspect of this disclosure provides a method for generating a trajectory planning model, comprising:

[0015] Obtain a training sample set, wherein the training sample includes at least one frame of input data and at least one frame of label data, wherein the input data includes environmental perception data, agent state data, and interaction data, and the label data includes trajectory labels and decision labels;

[0016] Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model.

[0017] A fourth aspect of this disclosure provides a method for planning vehicle trajectories, comprising:

[0018] Acquire vehicle environmental perception data, vehicle status data, and interaction data;

[0019] Based on the vehicle status data and the interaction data, determine the text command;

[0020] Based on the first sub-model in the trajectory planning model, the environmental perception data and the text instructions are processed to obtain the first multimodal feature;

[0021] Based on the first multimodal features and decision information features, the decision features are determined;

[0022] Based on the second sub-model in the trajectory planning model, the decision features, the environmental perception data, and the vehicle state data are processed to obtain the vehicle's driving trajectory in the future time period.

[0023] A fifth aspect of this disclosure provides a computer-readable storage medium storing computer program instructions that, when executed, implement the above-described vehicle trajectory planning method and / or the above-described trajectory planning model generation method.

[0024] A sixth aspect of this disclosure provides a computer program product including computer program instructions that, when executed by a processor, implement the above-described vehicle trajectory planning method and / or the above-described trajectory planning model generation method.

[0025] Based on the embodiments of this disclosure, when training the trajectory planning model, a training sample set can be obtained first. Each training sample includes at least one frame of input data and at least one frame of label data. The input data may include environmental perception data, agent state data, and interaction data. The label data includes trajectory labels and decision labels. Then, based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model. Therefore, by training the initial trajectory planning model in multiple stages using training samples from multiple dimensions, the embodiments of this disclosure can solve the problem of inconsistency between the decisions generated by the visual language model and the trajectory planned by the end-to-end model in trajectory planning models in related technologies, thereby improving the robustness and accuracy of trajectory planning using the trajectory planning model.

[0026] Based on the embodiments of this disclosure, when vehicle trajectory prediction is required, environmental perception data, vehicle state data, and interaction data of the vehicle can be acquired first. Based on the vehicle state data and interaction data, text commands are determined. Then, based on the first sub-model (visual language model) in the trajectory planning model, the environmental perception data and text commands are processed to obtain multimodal features. Next, based on the multimodal features and decision information, decision features are determined. Finally, based on the second sub-model (end-to-end model) in the trajectory planning model, the decision features, environmental perception data, and the vehicle state data are processed to obtain the vehicle's driving trajectory within a future time period. Therefore, in this embodiment of the disclosure, the multimodal features output by the visual language model can provide fine-grained semantic information, while the decision information features corresponding to the decision category provide decision category information. By superimposing and fusing the multimodal features and decision information features and inputting them into the end-to-end model, both semantic information and decision category information can be provided to the end-to-end model simultaneously. This helps improve the accuracy of trajectory planning by the end-to-end model, provides accurate driving control basis for intelligent driving vehicles, and enhances the driving safety of intelligent driving vehicles. Attached Figure Description

[0027] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0028] Figure 1 This is an application scenario diagram provided by an exemplary embodiment of this disclosure;

[0029] Figure 2 This is a structural diagram of a trajectory planning model generation apparatus and / or a vehicle trajectory planning apparatus provided in an exemplary embodiment of this disclosure;

[0030] Figure 3 This is an exemplary schematic diagram illustrating the generation of a trajectory planning model in an embodiment of this disclosure;

[0031] Figure 4 This is a schematic diagram of the framework for generating the trajectory planning model in an embodiment of this disclosure;

[0032] Figure 5 This is a flowchart illustrating step 302 provided in an exemplary embodiment of this disclosure;

[0033] Figure 6 This is a flowchart illustrating step 322 provided in an exemplary embodiment of this disclosure;

[0034] Figure 7 This is a flowchart illustrating step 3221 provided in an exemplary embodiment of this disclosure;

[0035] Figure 8 This is a flowchart illustrating step 3222 provided in an exemplary embodiment of this disclosure;

[0036] Figure 9 This is a flowchart illustrating step 323 provided in an exemplary embodiment of this disclosure;

[0037] Figure 10 This is a flowchart illustrating step 3231 provided in an exemplary embodiment of this disclosure;

[0038] Figure 11 This is a flowchart illustrating step 3232 provided in an exemplary embodiment of this disclosure;

[0039] Figure 12 This is an exemplary schematic diagram of vehicle trajectory planning in an embodiment of this disclosure;

[0040] Figure 13 This is a flowchart illustrating step 1204 provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0041] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0042] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0043] This disclosure outlines

[0044] In realizing this disclosure, the inventors discovered through research that in the field of intelligent driving (including assisted driving and autonomous driving), vehicles are typically equipped with multiple sensors to observe the vehicle's driving environment and its own state (pose, speed, yaw angle, etc.), and then plan the driving trajectory based on the observed data. Accurately planning the vehicle's driving trajectory in the future time period (e.g., the next 6 seconds) helps provide accurate driving control basis for intelligent driving vehicles and improves the driving safety of autonomous vehicles.

[0045] In related technologies, a scheme combining a Visual Language Model (VLM) with an End-to-End (E2E) model is used for trajectory planning to obtain the vehicle's driving trajectory over a future time period. However, due to the lack of a unified behavioral semantic alignment mechanism between the visual language model and the end-to-end model, a mismatch easily arises in real traffic environments between the decisions generated by the visual language model and the planned trajectories generated by the end-to-end model, resulting in low robustness and accuracy of vehicle trajectory prediction. For example, the decision obtained by the VLM model based on perception data and interaction data is "decelerate," while the trajectory planned by the E2E model based on perception data and vehicle state data reflects the decision of "changing lanes and overtaking," thus creating a mismatch.

[0046] In this embodiment of the disclosure, when training the trajectory planning model, the initial trajectory planning model can be trained in multiple stages. Through multiple stages of training, consistency constraints can be imposed on the output of the VLM model and the output of the E2E model. Furthermore, closed-loop interactive feedback is introduced during the training process, which helps to further reduce the deviation between the decision output of the VLM model and the planned trajectory output of the E2E model in the long time series, thereby improving the robustness and accuracy of trajectory planning using the trained trajectory planning model.

[0047] Exemplary Applications

[0048] The embodiments disclosed herein can be used for trajectory planning of autonomous mobile devices (also known as intelligent agents) such as vehicles, robots, and drones during driving.

[0049] Figure 1 This is an application scenario diagram provided by an exemplary embodiment of this disclosure, such as... Figure 1As shown, a computing platform 120 and multiple sensors 130 can be deployed on the vehicle 110 as needed. The sensors 130 may include sensors for collecting environmental perception data during vehicle operation and sensors for collecting their own state data.

[0050] Sensors 130, used to collect environmental perception data during vehicle operation, can be deployed at different locations and / or in different orientations on the vehicle 110 according to perception requirements, to perform environmental perception on corresponding areas of the external environment of the vehicle 110 (e.g., the left front area, the front area, the right front area, the left rear area, the rear area, the right rear area, the left side of the vehicle, the rear side of the vehicle, etc.). The at least one sensor 130 can include sensors of the same or different types, such as visual sensors (i.e., cameras), LiDAR (Light Detection and Ranging), millimeter-wave radar (MMW), ultrasonic sensor systems (USS), etc. During vehicle operation, the sensors 130 on the vehicle 110 can observe the surrounding environment of the vehicle 110 to obtain environmental perception data. This environmental perception data can include point cloud data and / or image data of the surrounding environment.

[0051] Sensors 130, used to collect their own state data, can be deployed at different locations within the vehicle 110 according to sensing needs, in order to sense the state of the vehicle 110. The vehicle's state may include vehicle pose, longitudinal acceleration, lateral acceleration, yaw rate, steering wheel angular rate, etc. During vehicle operation, the sensors 130 on the vehicle 110 can observe the state of the vehicle 110 and obtain its own state data.

[0052] The computing platform 120 includes a processor and can transmit data with the sensor 130 via a serial data bus or a controller area network (CAN) bus.

[0053] The multiple sensors 130 may include visual sensors (i.e., cameras), LiDAR (Light Detection and Ranging), millimeter-wave radar (MMW), and ultrasonic sensor systems (USS). Specifically, point cloud data of the environment surrounding the vehicle 110 can be obtained through LiDAR observation, and image data of the environment surrounding the vehicle 110 can be acquired through visual sensors.

[0054] During the driving process, vehicle 110 can acquire environmental perception data, vehicle status data, and interaction data. Based on the vehicle status data and interaction data, it determines text commands. Then, based on the first sub-model (visual language model) in the trajectory planning model, it processes the environmental perception data and text commands to obtain multimodal features. Based on the multimodal features and decision information, it determines decision features. Finally, based on the second sub-model (end-to-end model) in the trajectory planning model, it processes the decision features, environmental perception data, and vehicle status data to obtain the vehicle's driving trajectory in the future time period.

[0055] The trajectory planning model described above can be obtained by training the initial trajectory planning model through multiple stages using the acquired training sample set. The training sample set includes at least one frame of input data and at least one frame of label data. The input data in each frame includes environmental perception data, agent state data, and interaction data, while the label data includes trajectory labels and decision labels.

[0056] The training samples can be expert data. In some implementations, a skilled safety driver (expert) can drive the data collection vehicle in a real-world environment, collecting and storing environmental perception data, status data, and interaction data through sensors deployed on the vehicle, such as cameras, lidar, and inertial measurement units. The collected data is then labeled and verified by experts to obtain training samples. In other implementations, during road testing, a safety driver, in conjunction with an intelligent driving device, can drive the data collection vehicle in a real-world environment, collecting and storing environmental perception data, status data, and interaction data through sensors deployed on the vehicle, such as cameras, lidar, and inertial measurement units. The collected data is then labeled and verified by experts to obtain training samples.

[0057] In some embodiments, during data annotation, the driving behavior of the data collection vehicle during its journey can be annotated to obtain decision labels, and the vehicle's driving trajectory deemed safe and reasonable by experts can be determined as trajectory labels. In other embodiments, the vehicle's driving trajectory deemed safe and reasonable by experts can be determined as trajectory labels, and based on kinematic principles, the driving trajectory corresponding to the trajectory label can be parsed to determine the decision label. For example, the vehicle state data of two adjacent frames can be determined based on the driving trajectory, and the decision label corresponding to each frame can be determined based on this.

[0058] Figure 1 This is merely an exemplary application scenario implementation of the present disclosure. Those skilled in the art can understand from the description of the present disclosure that the present disclosure can also adopt any other feasible implementation. For example, the computing platform can be deployed in whole or in part on a cloud server or terminal device (such as a mobile terminal, tablet computer, PC, etc.), and receive sensor data collected by multiple sensors 130 through a communication connection with the vehicle driving control system. The data is then processed using the vehicle trajectory planning device provided in the present disclosure to obtain the vehicle's driving trajectory in the future time period and feed it back to the vehicle driving control device.

[0059] Exemplary device

[0060] Figure 2 This is a structural diagram of a trajectory planning model generation apparatus and / or a vehicle trajectory planning apparatus provided in an exemplary embodiment of this disclosure. The trajectory planning model generation apparatus and / or vehicle trajectory planning apparatus of this embodiment can be applied to in-vehicle electronic devices, specifically deployed on the computing platform of the in-vehicle electronic device, for example... Figure 1 The computing platform 120 of the vehicle 110 shown is used for this vehicle.

[0061] like Figure 1 As shown, the trajectory planning model generation device and / or vehicle trajectory planning device may include at least one processor 21. The processor 21 may be a central processing unit (CPU) or other processing unit (e.g., a SoC) with data processing capabilities and / or instruction execution capabilities. The processor 21 may be configured for generating trajectory planning models and / or planning vehicle trajectories.

[0062] Figure 3 This is an exemplary schematic diagram illustrating the generation of a trajectory planning model implemented by processor 21 in this embodiment of the present disclosure. See also Figure 3 The processor 21 can be configured as follows:

[0063] Obtain a training sample set, which includes at least one frame of input data and at least one frame of label data. The input data includes environmental perception data, agent state data, and interaction data. The label data includes trajectory labels and decision labels (301).

[0064] Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model (302).

[0065] At least one frame of input data in the training samples can be expert data.

[0066] In some implementations, a skilled safety driver (expert) can drive the data collection vehicle in a real-world environment, collecting and storing environmental perception data, status data, and interaction data through sensors deployed on the vehicle, such as cameras, lidar, and inertial measurement units. Training samples are then obtained through data annotation and expert verification. In other implementations, during road testing, a safety driver, in conjunction with an intelligent driving device, can drive the data collection vehicle in a real-world environment, collecting and storing environmental perception data, status data, and interaction data through sensors deployed on the vehicle, such as cameras, lidar, and inertial measurement units. Training samples are then obtained after data annotation and expert verification of the collected data.

[0067] In some embodiments, data annotation can be performed on the driving behavior of the data collection vehicle during its journey. For example, if deceleration or lane changing is recorded during the journey, corresponding deceleration and lane changing labels can be assigned to obtain decision labels. The vehicle's trajectory, deemed safe and reasonable by experts, can then be identified as the trajectory label. In other embodiments, the vehicle's trajectory, deemed safe and reasonable by experts, can be identified as the trajectory label. Based on kinematic principles, the agent's state data for each frame can be analyzed to determine the decision label.

[0068] The decision can be a lateral control decision or a longitudinal control decision, such as going straight, turning, accelerating, or decelerating.

[0069] In some embodiments, environmental perception data may include image data and point cloud data. During road testing, environmental perception data can be collected at a preset frame rate by any number of sensors of the same or different types deployed at different locations and / or for different orientations on an intelligent agent (such as a vehicle), such as visual sensors, lidar, millimeter-wave radar, ultrasonic radar, etc.

[0070] In some embodiments, the environmental perception data is purely visual data, acquired by at least one visual sensor, which may include a forward-looking main camera, a forward-looking wide-angle camera, a surround-view camera system, etc.

[0071] In some implementations, at least one sensor is deployed on the intelligent agent. The sensors are located and / or oriented differently. The sensors may include visual sensors and / or lidar. When there are multiple sensors, the environmental perception data collected by different sensors (of the same modality or different modalities) can be stitched together or aligned to ensure the consistency of the environmental perception data collected by different sensors.

[0072] In this embodiment, the agent's state data may include, but is not limited to, the agent's pose, driving direction, driving speed, driving acceleration, driving angular acceleration, yaw rate, and steering wheel angular rate. During road testing, agent state data can be collected using any number and type of sensors deployed at different locations on the agent (e.g., a vehicle), such as yaw rate sensors, acceleration sensors, velocity sensors, and inertial measurement units (IMUs). The various velocities and / or positions in the agent state data can be obtained in a specified coordinate system, which may include, but is not limited to, the vehicle coordinate system, the world coordinate system, or any one or more coordinate systems. These coordinate systems can also be converted to each other based on the relationships between them, thereby obtaining agent state data in the same coordinate system.

[0073] In this embodiment of the disclosure, the interactive data can be data input to the trajectory planning model generation device in different ways, including but not limited to instruction data for controlling the intelligent agent's driving input by the user through natural voice, touch screen, etc., navigation data for controlling the intelligent agent's driving input by the associated application, and pre-set template-type interactive data for generating text instructions. For example, a user can input the instruction "Navigate to the nearest gas station" via voice, and the template-type interactive data can be "How to drive in this scenario," with the navigation application outputting navigation data such as "Drive along the current road, and enter the main road after a series of consecutive right turns." In some embodiments, when user interaction data or navigation data input by the user's voice is obtained, the user interaction data and / or navigation data can be parsed to extract driving-related keywords, such as "turn left" or "enter the main road," and these keywords can be added to the text instruction template to generate text instructions.

[0074] Among them, the trajectory label is the actual driving trajectory of the vehicle during the road test, and the decision label is the actual decision made by the vehicle during the road test.

[0075] In this embodiment of the disclosure, the multi-frame input data in the training samples are multi-frame data with temporal relationships collected by the acquisition vehicle.

[0076] When acquiring trajectory labels, the trajectory labels corresponding to each frame of input data can be divided from a long real driving trajectory. The trajectory label corresponding to each frame of input data is a segment of trajectory starting from the corresponding frame of data.

[0077] When acquiring decision labels, the decision label corresponding to each frame of input data can be determined based on the agent's state data in each frame of input data. For example, the acceleration in the current frame can be used to determine whether it is an acceleration or deceleration decision, and the lateral acceleration in the current frame can be used to determine whether it is a left turn or right turn decision.

[0078] For example, driving data for 30 seconds was collected at a frame rate of 30 frames per second. Therefore, the actual driving trajectory is a 30-second, 900-frame driving trajectory. The trajectory label for each frame of input data is the trajectory for 3 seconds. Thus, the trajectory label for the first frame of input data is the trajectory for the 3 seconds and 90 frames following the first frame of input data. If the acceleration in the state data of the first frame of input data is greater than the set acceleration threshold, then the decision label corresponding to the first frame of input data can be determined as acceleration.

[0079] In this embodiment, the initial trajectory planning model is a model comprising a first sub-model and a second sub-model. The first sub-model can be a multimodal model, such as a visual-language model. In this embodiment, the visual-language model can achieve cross-modal understanding and generate predictive decisions (such as acceleration, deceleration, right turn, left turn) based on a combination of environmental perception data and text commands. The second sub-model can be an end-to-end model, and in this embodiment, the second sub-model can be used to output a predicted trajectory based on the input training samples.

[0080] The multi-stage training method for the initial trajectory planning model in this embodiment can be found in subsequent embodiments, which will not be detailed here.

[0081] The aforementioned trajectory planning model generation apparatus and / or vehicle trajectory planning apparatus may further include at least one memory 22. The memory 22 may store one or more computer program products and may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 21 may execute one or more computer program instructions to implement the trajectory planning model generation, vehicle trajectory planning, and / or other desired functions of the various embodiments of this disclosure.

[0082] In one example, the trajectory planning model generation device and / or vehicle trajectory planning device may further include an input device 23 and an output device 24, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0083] The input device 23 may also include, for example, a keyboard, a mouse, etc.

[0084] The output device 24 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0085] Of course, for the sake of simplicity, Figure 2 This document only shows some of the components in the trajectory planning model generation apparatus that are relevant to this disclosure, omitting components such as buses, input / output interfaces, etc. In addition, the generation of the trajectory planning model may include any other suitable components depending on the specific application.

[0086] It should be noted that the trajectory planning model in this disclosed technical solution can be set on various types of electronic devices, including but not limited to mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.

[0087] In some alternative implementations, see [link to relevant documentation]. Figure 4 The diagram illustrates a multi-stage training framework for a trajectory planning model.

[0088] First, the first and second sub-models in the initial trajectory planning model are trained separately. Specifically, the first sub-model is trained to obtain the third sub-model by using environmental perception data and text instructions (generated based on interaction data and agent state data) and decision labels in the training samples, and the second sub-model is trained to obtain the fourth sub-model by using environmental perception data, agent state data and trajectory labels in the training samples.

[0089] Secondly, the initial trajectory planning model is trained through joint training. Specifically, joint training can be implemented in two processes. In the first process, the third and fourth sub-models are connected through an initial connection layer, and the third sub-model is set to a frozen state. Environmental perception data and text commands from the training samples are input into the third sub-model, which outputs multimodal features and a second prediction decision. These output features and the second prediction decision are then input into the initial connection layer for processing, where the initial connection layer outputs decision features. These decision features, along with environmental perception data and agent state data from the training samples, are then input into the fourth sub-model for processing, where the fourth sub-model outputs the second predicted trajectory. Based on the second predicted trajectory and trajectory labels, the initial connection layer and the fourth sub-model are adjusted. The parameters are used to obtain a second trajectory planning model (including a third sub-model, a first connecting layer obtained by adjusting the parameters of the initial connecting layer, and a fifth sub-model obtained by adjusting the parameters of the fourth sub-model). In the second process of joint training, the environmental perception data and text commands in the training samples are first input into the third sub-model, which outputs multimodal features and a third prediction decision. The output multimodal features and the second prediction decision are then input into the first connecting layer for processing, and the first connecting layer outputs decision features. These decision features, along with the environmental perception data and agent state data in the training samples, are then input into the fifth sub-model for processing, and the fifth sub-model outputs a third prediction trajectory. Kinematic constraints are used to determine the matching degree between the third prediction decision and the third prediction trajectory, and selective optimization training is performed using inconsistent (matching) samples. This training process alleviates the deviation between the third and fifth sub-models in the prediction decision-predicted trajectory mapping, achieves explicit consistency constraints, and obtains a first trajectory planning model (including a sixth sub-model obtained by adjusting the parameters of the third sub-model, a second connecting layer obtained by adjusting the parameters of the first connecting layer, and a seventh sub-model obtained by adjusting the parameters of the fifth sub-model).

[0090] Finally, the first trajectory planning model is further trained using closed-loop training. During closed-loop training, the first frame of input data from the training samples is input into the first trajectory planning model, which outputs the fourth predicted trajectory and the first predicted decision. Then, a 3D reconstruction model is used to process the time and pose information corresponding to the first frame data in the fourth predicted trajectory, rendering a virtual scene image. The virtual scene image, the agent state data corresponding to the first frame data in the fourth predicted trajectory, and text commands (determined based on the agent state data and interaction data corresponding to the first frame data in the fourth predicted trajectory) are repeatedly input into the first trajectory planning model to achieve interactive iterative prediction. The system obtains a first predicted trajectory of a preset duration (based on the pose information of the first frame data in each output fourth predicted trajectory) and multiple first predicted decisions. Then, it uses real-world environment data determined by at least one frame of input data from the training samples to perform collision risk detection on the first predicted trajectory. If it is determined that there is a collision risk when the agent follows the first predicted trajectory, the first predicted trajectory is optimized. Based on the optimized predicted trajectory, the first predicted decisions are further optimized to generate optimized predicted decisions. Finally, the first trajectory planning model is trained using the first predicted trajectory, first predicted decisions, optimized predicted trajectory, and optimized predicted decisions to obtain the trajectory planning model. The aforementioned 3D reconstruction model is trained based on at least one frame of input data from the training samples. It is trained by using the first frame of input data from the training samples as input and other frames that are temporally subsequent as ground truth. This 3D reconstruction model can render a virtual scene image based on the time and pose information corresponding to the first frame of input data. This training process trains the model through closed-loop interactive feedback with the 3D reconstruction model. This helps to significantly improve the performance of the trajectory planning model through real-time data iteration and dynamic optimization, reduce the bias between the decision and trajectory of the trained trajectory planning model, and improve the robustness, long-term stability and trajectory prediction accuracy of the trajectory planning model in dynamic scenarios.

[0091] In some alternative implementations, see [link to relevant documentation]. Figure 5 In section 302 above, the specific process by which processor 21 performs multi-stage training on the initial trajectory planning model based on the training sample set to obtain the trajectory planning model may include:

[0092] Based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models (321);

[0093] Based on the training sample set, the third sub-model and the fourth sub-model are jointly trained to obtain the first trajectory planning model (322);

[0094] Based on the training sample set, the first trajectory planning model is trained to obtain the trajectory planning model (323).

[0095] In this embodiment, the first sub-model can be a visual language model. First, text instructions for inputting the first sub-model can be determined based on agent state data and interaction data in the training samples. Then, based on environmental perception data, text instructions, and decision labels, the first sub-model in the initial trajectory planning model is trained to obtain the third sub-model.

[0096] In practice, text instructions can be generated using pre-defined instruction templates. For example, adding agent state data and interaction data to the corresponding positions in the instruction template will produce text instructions. For instance, the instruction template for the first task instruction is: "Given the given image, interaction data, and agent state data, where the state data is [**] and the interaction data is [**], please provide a decision on how to drive in this scenario."

[0097] For example, if the agent's state data is "vehicle speed 50 km / h, engine speed 600 rpm, steering wheel angle 10 degrees" and the interaction data is "turn left at the intersection ahead", then according to the above template, the text instruction can be generated as follows: "Given the given image, interaction data, and agent state data, with the state data being [vehicle speed 50 km / h, engine speed 600 rpm, steering wheel angle 10 degrees] and the interaction data being [turn left at the intersection ahead], please provide a decision on how to drive in this scenario."

[0098] After generating the text instructions, the text instructions and environmental perception data can be input into the first sub-model. The first sub-model processes the text instructions and environmental perception data, outputs a prediction decision, and then determines the model loss value based on the prediction decision and the decision label, and trains the first sub-model.

[0099] In this embodiment of the disclosure, the second sub-model can be an end-to-end model. The fourth sub-model can be obtained by training the second sub-model in the initial trajectory planning model based on environmental perception data, agent state data, and trajectory labels.

[0100] In practice, environmental perception data and agent state data are input into the second sub-model. The second sub-model processes the environmental perception data and agent state data, outputs the predicted trajectory, and then determines the model loss value based on the predicted trajectory and trajectory label, and trains the second sub-model.

[0101] The training processes of the first and second sub-models are both optimization processes. The optimal solution fitting process for the first and second sub-models is mainly carried out iteratively by minimizing the error. For an input training sample, the decision loss function value can be calculated using the output prediction decision and the corresponding decision label. This decision loss function value is propagated to the connection between each neuron in the first sub-model through the backpropagation algorithm. Then, the parameters of the first sub-model are updated and modified using the gradient descent algorithm, so that the decision loss value calculated during the iterative training process gradually decreases until the preset training completion condition is met, and the third sub-model can be obtained from the first sub-model. Similarly, the trajectory loss function value can be calculated using the output prediction trajectory and the corresponding trajectory label. This trajectory loss function value is propagated to the connection between each neuron in the second sub-model through the backpropagation algorithm. Then, the parameters of the second sub-model are updated and modified using the gradient descent algorithm, so that the trajectory loss value calculated during the iterative training process gradually decreases until the preset training completion condition is met, and the fourth sub-model can be obtained from the second sub-model.

[0102] In this embodiment of the disclosure, when determining the loss function value using the prediction decision and the corresponding decision label, the loss function corresponding to the classification task, such as the cross-entropy loss function or the log loss function, can be used to determine the loss function value; when determining the loss function value using the prediction trajectory and trajectory label, the loss function corresponding to the regression task, such as the mean squared error loss function or the mean absolute error loss function, can be used to determine the loss function value.

[0103] The preset training completion conditions may include, for example, at least one of the following: the difference between the predicted decision and the decision label (or, for the second sub-model, the difference between the predicted trajectory and the trajectory label) is less than a first preset threshold; the number of iterative training iterations for the first sub-model (or the second sub-model) reaches a preset number (e.g., 10,000 times); etc. The embodiments of this disclosure do not limit the preset training completion conditions.

[0104] In this embodiment, the joint training of the third and fourth sub-models to obtain the first trajectory planning model refers to a training method in which the third and fourth sub-models optimize the model during training through parameter sharing or weighted summation, ultimately outputting a comprehensive result. For specific implementation details, please refer to [link to relevant documentation]. Figure 6 The illustrated embodiment.

[0105] In this embodiment of the disclosure, the first trajectory planning model is trained based on the training sample set. For a detailed implementation of the trajectory planning model, please refer to [link to relevant documentation]. Figure 9 The illustrated embodiment.

[0106] Based on the embodiments of this disclosure, a multi-stage model training method is disclosed. First, separate model training helps to allow different models to optimize independently and improve the model training rate. Then, joint training helps to achieve consistency correction of the decisions and trajectories output by the two models, continuously improve the synergy of the two models, and avoid misjudgment, late judgment or dangerous behavior caused by decision-trajectory deviation.

[0107] In some alternative implementations, see [link to relevant documentation]. Figure 6 In section 322 above, the specific process by which processor 21 performs multi-stage training on the initial trajectory planning model based on the training sample set to obtain the trajectory planning model may include:

[0108] Based on the training sample set, the initial trajectory planning model is trained to update the model parameters in the initial trajectory planning model except for the third sub-model, and the second trajectory planning model (3221) is obtained.

[0109] Based on the training sample set, the second trajectory planning model is trained to obtain the first trajectory planning model (3222).

[0110] In this embodiment, the initial trajectory planning model includes a third sub-model, a fourth sub-model, and an initial connection layer, with the third and fourth sub-models connected through the initial connection layer. After training the initial trajectory planning model using step 3221, the third sub-model can be frozen first. Then, training samples are input into this model, which includes the third sub-model, the initial connection layer, and the fourth sub-model, to train the initial trajectory model, thereby updating the model parameters in the initial trajectory planning model excluding the third sub-model, thus obtaining the second trajectory planning model.

[0111] The initial connection layer is a connector that connects the third sub-model (visual language model) and the fourth sub-model (end-to-end model). It is used to encode the decision output by the third sub-model to obtain decision information features, and to fuse the decision information features and the multimodal features output by the third sub-model to obtain decision features. The fused decision features are then input into the fourth sub-model for trajectory prediction.

[0112] In the process of fusing decision information features and multimodal features output by the third sub-model, the decision information features and multimodal features can be concatenated first, and then the concatenated features can be fused into a single decision feature through a multilayer perceptron (MLP); or, the decision information features and multimodal features can be added first, and then the concatenated features can be fused into a single decision feature through a multilayer perceptron (MLP).

[0113] In some alternative implementations, see [link to relevant documentation]. Figure 7In section 3221 above, the specific process by which processor 21 trains the initial trajectory planning model based on the training sample set to update the model parameters in the initial trajectory planning model excluding the third sub-model, and obtains the second trajectory planning model, may include:

[0114] Based on the third sub-model in the initial trajectory planning model, the environmental perception data and text instructions in the training samples are processed to obtain multimodal features and the second prediction decision (32211).

[0115] Based on the initial connection layer in the initial trajectory planning model, the multimodal features and the second prediction decision are processed to obtain the decision features (32212);

[0116] Based on the fourth sub-model in the initial trajectory planning model, the decision features and environmental perception data and agent state data in the training samples are processed to obtain the second predicted trajectory (32213).

[0117] Based on the second predicted trajectory and trajectory label, the initial connection layer and the fourth sub-model in the initial trajectory planning model are trained to obtain the second trajectory planning model (32214).

[0118] Among them, multimodal features are unified semantic representations generated by the visual language model through cross-modal alignment of input environmental perception data and text instructions. Decoding multimodal features can yield a second prediction decision, such as "accelerate straight ahead".

[0119] In this embodiment of the disclosure, after the initial connection layer outputs decision features, the decision features and environmental perception data and agent state data in the training samples can be input into the fourth sub-model for processing. The fourth sub-model outputs the second predicted trajectory, and the parameters of the initial connection layer and the fourth sub-model are adjusted based on the second predicted trajectory and trajectory label to obtain the second trajectory planning model.

[0120] The second trajectory planning model may include a third sub-model, a first connection layer obtained by adjusting the parameters of the initial connection layer, and a fifth sub-model obtained by adjusting the parameters of the fourth sub-model.

[0121] In this embodiment, the training process of the initial connection layer and the fourth sub-model in the initial trajectory planning model is an optimal solution solving process. For an input training sample, the trajectory loss function value can be calculated using the output second predicted trajectory and trajectory label. The decision loss function value is propagated to the connection between each neuron in the initial connection layer and the fourth sub-model through the backpropagation algorithm. Then, the parameters of the initial connection layer and the fourth sub-model are updated and modified using the gradient descent algorithm, so that the trajectory loss value calculated during the iterative training process gradually decreases until the preset training completion condition is met. Thus, the first connection layer is obtained from the initial connection layer, and the fifth sub-model is obtained by adjusting the parameters of the fourth sub-model, resulting in the second trajectory planning model.

[0122] In this embodiment of the disclosure, when determining the loss function value using the second predicted trajectory and trajectory label, the loss function corresponding to the regression task, such as the mean squared error loss function or the mean absolute error loss function, can be used to determine the loss function value.

[0123] In some alternative implementations, see [link to relevant documentation]. Figure 8 In the above 3222, the specific process by which the processor 21 trains the second trajectory planning model based on the training sample set to obtain the first trajectory planning model may include:

[0124] Based on the second trajectory planning model, the training samples are processed to obtain the third predicted trajectory and the third predicted decision (32221);

[0125] Based on kinematic constraints, the matching degree between the third predicted trajectory and the third predicted decision is determined (32222);

[0126] In response to the mismatch between the matching degree characterization of the third predicted trajectory and the third predicted decision, the second trajectory planning model is trained based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model (32223).

[0127] In this embodiment of the disclosure, the third sub-model in the second trajectory planning model can be used to process the environmental perception data and text commands to obtain the third prediction decision and multimodal features. Then, the third prediction decision is encoded through the first connection layer to obtain decision information features. The decision information features and multimodal features are then fused to obtain decision features. The fifth sub-model in the second trajectory planning model processes the decision features input from the first connection layer and the environmental perception data and agent state data in the training samples, and the fifth sub-model outputs the third prediction trajectory.

[0128] In the process of fusing decision information features and multimodal features, the decision information features and multimodal features can be concatenated first, and then the concatenated features can be fused into a single decision feature through MLP; or, the decision information features and multimodal features can be added first, and then the concatenated features can be fused into a single decision feature through multilayer perceptron MLP.

[0129] For any training sample, the matching degree between the third prediction trajectory and the third prediction decision is determined using kinematic constraints.

[0130] In some implementations, kinematic constraints can be constructed based on a vehicle kinematics model, and the third prediction decision and agent state data can be processed according to the kinematic constraints to predict the agent's future position in the next frame. Then, the difference between the predicted future position in the next frame and the position in the next frame corresponding to the third predicted trajectory is calculated (e.g., Euclidean distance is calculated), and the matching degree is determined based on the difference.

[0131] In other implementations, kinematic constraints can be constructed based on the vehicle kinematics model, and the third predicted trajectory can be processed according to the kinematic constraints to determine the agent state data. Based on the agent state data, the decision corresponding to the current frame can be determined, and then it can be determined whether the decision type determined according to the third predicted trajectory matches the decision type corresponding to the third predicted decision.

[0132] For example, if the decision for the current frame determined based on the agent's state data from the third predicted trajectory is to accelerate straight ahead, while the third predicted decision is to decelerate or turn, then it can be determined that the third predicted trajectory and the third predicted decision do not match.

[0133] In this embodiment of the disclosure, after determining the training samples where the third predicted trajectory and the third predicted decision do not match based on the matching degree, the loss function value can be determined based on the mismatched third predicted trajectory and the third predicted decision, as well as the trajectory label and the decision label, and the model parameters of the second trajectory planning model can be updated according to the loss function value to achieve training; alternatively, the training samples where the third predicted trajectory and the third predicted decision do not match can be used as training samples for the second trajectory planning model to train the second trajectory planning model.

[0134] In some optional implementations, the specific process of training the second trajectory planning model to obtain the first trajectory planning model based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label may include: determining the first loss function value based on the third predicted decision and the decision label; determining the second loss function value based on the third predicted trajectory and the trajectory label; performing a weighted operation on the first loss function value and the second loss function value to obtain the third loss function value; and training the second trajectory planning model based on the third loss function value to obtain the first trajectory planning model.

[0135] The training process for the second trajectory planning model is essentially an optimal solution-finding process, primarily achieved through iterative error minimization. For training samples with inconsistent decision-trajectory relationships, the first loss function value can be calculated using the corresponding third predicted decision and decision label, while the second loss function value can be calculated using the corresponding third trajectory decision and trajectory label. A weighted average of the first and second loss function values ​​yields the third loss function value. This third loss function value is then propagated to the connections between each neuron in the second trajectory planning model via backpropagation. Gradient descent is then used to update and modify the parameters of the second trajectory planning model, gradually decreasing the calculated third loss function value during iterative training until the preset training completion conditions are met. This process allows the second trajectory planning model to be used to derive the first trajectory planning model.

[0136] Specifically, when determining the first loss function value using the third prediction decision and decision label, the loss function corresponding to the classification task, such as the cross-entropy loss function and the log loss function, can be used to determine the loss function value; when determining the second loss function value using the third trajectory decision and trajectory label, the loss function corresponding to the regression task, such as the mean squared error loss function and the mean absolute error loss function, can be used to determine the loss function value.

[0137] Based on the embodiments of this disclosure, by determining the matching degree between the predicted decision and the predicted trajectory, training samples with inconsistent decisions and trajectories are identified, and the second trajectory planning model is further trained using these training samples to achieve explicit consistency constraints. This enhances the control of the decision output by the visual language model in the second trajectory planning model over the trajectory planning of the end-to-end model, and helps to promote the synchronization of semantics and planning trends between the visual language model and the end-to-end model in the trained first trajectory planning model.

[0138] In some alternative implementations, see [link to relevant documentation]. Figure 9 In section 323 above, the specific process by which processor 21 trains the first trajectory planning model based on the training sample set to obtain the trajectory planning model may include:

[0139] Based on the first trajectory planning model, the first frame of input data in at least one frame of input data is iteratively processed to obtain the first predicted trajectory and at least one first prediction decision (3231);

[0140] Based on the first prediction decision and the first prediction trajectory, the first trajectory planning model is trained to obtain the trajectory planning model (3232).

[0141] In this embodiment, the first trajectory planning model is further trained in a closed-loop environment. During the closed-loop training process, only the first frame input data from the training samples is input into the first trajectory planning model. This first frame input data can be denoted as the first frame input data. The first trajectory planning model then processes the first frame input data to output a predicted trajectory and prediction decisions, which can be denoted as the data corresponding to the first frame input data. The first frame data from the predicted trajectory is then used to interact with the 3D reconstruction model to obtain new input data for the first trajectory planning model, which can be denoted as the second frame input data. This new input data (i.e., the second frame input data) is then processed by the first trajectory planning model to obtain a new predicted trajectory and prediction decisions, which can be denoted as the data corresponding to the second frame input data. After multiple rounds of interaction with the 3D reconstruction model, a first predicted trajectory of a preset duration and multiple corresponding first prediction decisions can be obtained.

[0142] For details, see Figure 10 In the above 3231, the specific process by which the processor 21 iteratively processes the first frame of input data in at least one frame of input data based on the first trajectory planning model to obtain the first predicted trajectory and at least one first prediction decision may include:

[0143] Based on the first trajectory planning model, the first frame input data is processed to obtain the fourth predicted trajectory and the first prediction decision (32311) corresponding to the first frame input data;

[0144] Determine the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory (32312);

[0145] Based on the 3D reconstruction model, the pose and time information corresponding to the first frame data are processed to render and generate a virtual scene image (32313).

[0146] Based on the first trajectory planning model, the agent state data corresponding to the virtual scene image and the first frame data are processed to obtain the fourth predicted trajectory and the first predicted decision (32314) corresponding to the first frame data.

[0147] Repeatedly execute the operation of determining the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory to obtain the first predicted trajectory and at least one first prediction decision (32315).

[0148] In this embodiment of the disclosure, before reusing the first trajectory planning model to process the input data, the model is first set to a frozen state so that the model parameters of the first trajectory planning model remain unchanged during the interactive iteration process of determining the first predicted trajectory in the closed-loop environment.

[0149] In specific implementation, the environmental perception data and text commands (generated based on the interaction data and agent state data in the first frame input data) corresponding to the first frame input data can be processed by the sixth sub-model in the first trajectory planning model to obtain the first prediction decision and multimodal features. Then, the first prediction decision is encoded using the second connection layer to obtain decision information features. The decision information features and multimodal features are then fused to obtain decision features. The obtained decision features are input into the seventh sub-model, which processes the decision features, environmental perception data, and agent state data corresponding to the first frame input data to output the fourth prediction trajectory. Then, the pose information, time information, and agent state data corresponding to the first frame data in the fourth prediction trajectory are obtained. The pose information and time information are input into the 3D reconstruction model, which renders and outputs a virtual scene image. Based on the agent state data corresponding to the first frame data in the fourth prediction trajectory, a new text command is generated using the command template. Thus, the virtual scene image, agent state data, and new text command can be input into the first trajectory planning model for a new round of decision prediction and trajectory prediction. Repeating the above process can yield multiple fourth prediction trajectories and multiple first prediction decisions.

[0150] Among them, 32312 to 32314 can be an iterative execution process. Each execution of steps 32312 to 32314 can generate a first prediction decision and a fourth prediction trajectory. By iteratively executing steps 32312 to 32314, a first prediction trajectory that meets the preset conditions can be generated in a closed-loop environment.

[0151] It should be noted that in this embodiment, the predicted trajectory and predicted label obtained by processing the first frame input data and each subsequent frame input data obtained from the training data based on the first trajectory planning model are referred to as the fourth predicted trajectory and the first predicted label. These terms are only used to refer to the decisions and trajectories output by the first trajectory planning model. Different input data into the first trajectory planning model will result in different fourth predicted trajectories and potentially different first predicted decisions.

[0152] Among them, the first predicted trajectory that meets the preset conditions can be the one that reaches the preset length. Each iteration can obtain a fourth predicted trajectory, which is a trajectory point in the first predicted trajectory (the trajectory point corresponding to the first frame data in the fourth predicted trajectory). If the number of repeated executions reaches the preset number (e.g., 300 times), the number of trajectory points corresponding to the preset number can be obtained, which is the first predicted trajectory of the preset length.

[0153] The length of the first predicted trajectory can be the same as or less than the length of the actual driving trajectory corresponding to at least one frame of input data.

[0154] For example, if at least one frame of input data is environmental perception data, agent state data, and interaction data corresponding to a 20-second real driving trajectory, then the length of the first predicted trajectory can be the predicted trajectory corresponding to 20 seconds or a predicted trajectory of less than 20 seconds.

[0155] In this embodiment, the 3D reconstruction model is a model trained based on at least one frame of input data from the training samples. Specifically, the first frame of input data from the training samples can be used as input, and other frames that are sequentially located after the first frame of input data can be used as ground truth for training. The 3D reconstruction model can be a 3D Gaussian Splatting (3DGS) model, which can render a virtual scene image based on the input temporal and pose information.

[0156] In this embodiment of the disclosure, a method for training a 3D reconstruction model in a dynamic scene can be used to train the 3D reconstruction model. The environmental perception data in at least one frame of input data in the training samples is preprocessed. The Structure from Motion (SfM) algorithm is used to estimate the camera pose and generate a sparse point cloud from the image sequence corresponding to the environmental perception data. Then, the sparse point cloud is used as the initial scene representation and converted into a set of 3D Gaussian ellipsoids. Each Gaussian ellipsoid contains learnable attributes such as position, covariance matrix (describing shape and orientation), and opacity. After obtaining the 3D Gaussian ellipsoids, the 3D reconstruction model can be obtained by iterative training through an iterative training phase centered on differentiable rendering.

[0157] In some alternative implementations, see [link to relevant documentation]. Figure 11 In the above 3232, the specific process by which the processor 21 trains the first trajectory planning model based on the first prediction decision and the first prediction trajectory to obtain the trajectory planning model may include:

[0158] Based on the first predicted trajectory and the real environment data corresponding to at least one frame of input data, the collision risk value (32321) is determined.

[0159] In response to the collision risk value indicating the existence of collision risk, the first predicted trajectory and the first predicted decision are optimized to obtain the optimized predicted trajectory and the optimized predicted decision (32322);

[0160] Based on the first predicted trajectory, the first predicted decision, the optimized predicted trajectory, and the optimized predicted decision, the first trajectory planning model is trained to obtain the trajectory planning model (32323).

[0161] The real-world environment data corresponding to at least one frame of input data may include the motion trajectory of dynamic targets and the actual position information of static targets in a real driving environment. This real-world environment data corresponds to the environmental perception data in the training samples and is pre-labeled data.

[0162] Dynamic targets can include pedestrians, animals, vehicles, bicycles, motorcycles, etc., in a realistic driving environment. Static targets can include information such as curbs, lane lines, stop lines, zebra crossings, road arrows, virtual lane lines, elevated objects, guardrails, buildings, flowers, trees, etc.

[0163] In this embodiment of the disclosure, the collision risk value is used to indicate the risk of collision between the intelligent agent and dynamic and static targets in the driving environment. The collision risk can be information such as whether a collision will occur, whether a collision will not occur, or whether a collision is possible. Whether a collision with each dynamic target will occur can be calculated based on the intelligent agent's first predicted trajectory and the motion trajectory of the center point of each dynamic target. Whether a collision with each static target will occur can be calculated based on the intelligent agent's first predicted trajectory and the actual position information of each static target.

[0164] In this embodiment of the disclosure, the time-to-collision (TTC) time between the vehicle and each dynamic and / or static target can be calculated, and the obtained collision time can be compared with a preset collision time threshold. If the collision time is less than the collision time threshold, it can be determined that there is a risk of collision. If the collision time is not less than the collision time threshold, it can be determined that there is no risk of collision.

[0165] When it is determined that the agent may be at risk of colliding with any dynamic or static target, the first predicted trajectory can be optimized based on predetermined rules. For example, the length of the first predicted trajectory can be shortened, which can be achieved by slowing down the vehicle. For instance, the first predicted trajectory, which travels 6 meters per second, can be shortened to a trajectory that travels 5 meters per second.

[0166] In this embodiment of the disclosure, the predetermined rule can be a fixed rule, for example, the predetermined rule is to shorten the first predicted trajectory to 0.9 times; the predetermined rule can also be adaptively changed based on the severity of the collision risk. For example, if it is determined that the agent may collide with multiple targets or may collide with them in a very short time (e.g., 0.5 seconds) when traveling according to the first predicted trajectory, the first predicted trajectory can be shortened to 0.7 times; if it is determined that the agent may collide with one target or may collide with them in a slightly longer time (e.g., 2 seconds) when traveling according to the first predicted trajectory, the first predicted trajectory can be shortened to 0.9 times.

[0167] After determining the optimized predicted trajectory, the decision corresponding to each frame of data can be determined based on the optimized predicted trajectory. Specifically, the agent state data corresponding to each frame of data in the optimized predicted trajectory can be determined based on the agent state data corresponding to each frame of data in the first predicted trajectory. This agent state data corresponding to each frame of data includes velocity, pose, driving direction, driving speed, driving acceleration, driving angular acceleration, and yaw rate. Further, the decision for the corresponding frame, such as acceleration, deceleration, or maintaining speed, can be determined based on the velocity and position information in the agent state data corresponding to each frame of data in the optimized predicted trajectory.

[0168] For example, if the speed of the previous frame of data is less than the speed of the next frame of data in two adjacent frames, then the decision corresponding to the previous frame of data can be determined as deceleration.

[0169] In some optional implementations, the specific process of training the first trajectory planning model based on the first predicted trajectory, the first predicted decision, the optimized predicted trajectory, and the optimized predicted decision to obtain the trajectory planning model may include: determining the fourth loss function value based on the first predicted decision and the optimized predicted decision; determining the fifth loss function value based on the first predicted trajectory and the optimized predicted trajectory; performing a weighted operation on the fourth loss function value and the fifth loss function value to obtain the sixth loss function value; and training the first trajectory planning model based on the sixth loss function value to obtain the trajectory planning model.

[0170] Specifically, when determining the fourth loss function value using the first prediction decision and the optimized prediction decision, the loss function corresponding to the classification task, such as the cross-entropy loss function and the log loss function, can be used to determine the loss function value; when determining the fifth loss function value using the first prediction trajectory and the optimized prediction trajectory, the loss function corresponding to the regression task, such as the mean squared error loss function and the mean absolute error loss function, can be used to determine the loss function value.

[0171] The training process for the first trajectory planning model is an optimal solution-finding process, which is mainly achieved iteratively by minimizing the error. For the first frame of input data in the training samples, the fourth loss function value can be calculated using the first prediction decision and the optimized prediction decision. The fifth loss function value can be determined using the first predicted trajectory and the optimized predicted trajectory. Then, by weighting the fourth and fifth loss function values, the sixth loss function value can be obtained. The sixth loss function value is then propagated to the connections between each neuron in the first trajectory planning model through the backpropagation algorithm. The gradient descent algorithm is then used to update and modify the parameters of the first trajectory planning model, so that the sixth loss function value calculated during the iterative training gradually decreases until the preset training completion conditions are met. Thus, the trajectory planning model can be obtained from the first trajectory planning model.

[0172] Since there are multiple first prediction decisions and multiple corresponding optimal prediction decisions, the differences between these multiple first prediction decisions and multiple optimal prediction decisions can be determined to obtain the fourth loss function value. For example, by treating the multiple first prediction decisions as a vector and the multiple optimal prediction decisions as a vector, the difference between the two vectors can be calculated to obtain the fourth loss function value.

[0173] Based on the embodiments of this disclosure, a model is trained through closed-loop interactive feedback, which helps to significantly improve the performance of the trajectory planning model through real-time data iteration and dynamic optimization, reduce the deviation between decisions and trajectories of the trained trajectory planning model, and improve the robustness, long-term stability and trajectory prediction accuracy of the trajectory planning model in dynamic scenarios.

[0174] After generating the trajectory planning model through the above embodiments, vehicle trajectories can be planned based on the generated trajectory planning model. Figure 12 This is an exemplary schematic diagram of vehicle trajectory planning implemented by processor 21 in this embodiment of the present disclosure:

[0175] Acquire vehicle environmental perception data, vehicle status data, and interaction data (1201);

[0176] Based on vehicle status data and interaction data, determine the text command (1202);

[0177] Based on the first sub-model in the trajectory planning model, environmental perception data and text commands are processed to obtain multimodal features (1203);

[0178] Based on multimodal features and decision information, decision features are determined (1204);

[0179] Based on the second sub-model in the trajectory planning model, the decision features, environmental perception data and vehicle state data are processed to obtain the vehicle's driving trajectory (1205) in the future time period.

[0180] The future time period is defined as a continuous time interval after the current moment, such as within 6 seconds after the current moment.

[0181] In this embodiment, environmental perception data may include image data and point cloud data. It can be acquired by any number of sensors of the same or different types deployed at different locations on the vehicle and / or for different orientations, such as visual sensors, lidar, millimeter-wave radar, ultrasonic radar, etc., at a preset frame rate. Vehicle state data may include, but is not limited to, vehicle pose, driving direction, driving speed, driving acceleration, driving angular acceleration, yaw rate, steering wheel angular rate, etc. Interactive data can be data input to the trajectory planning model generation device in different ways, including, but not limited to, user input via natural voice, touchscreen, etc., for controlling the intelligent agent's driving, and navigation data input by associated applications for controlling the intelligent agent's driving.

[0182] In some embodiments, the first sub-model may be the visual language model in the trajectory planning model trained through the above embodiments, and the second sub-model may be the end-to-end model in the trajectory planning model trained through the above embodiments.

[0183] In some embodiments, the environmental perception data is purely visual data, acquired by at least one visual sensor, which may include a forward-looking main camera, a forward-looking wide-angle camera, a surround-view camera system, etc.

[0184] In some alternative implementations, after step 1203, processor 21 may also be configured to decode the first multimodal feature to obtain decision information.

[0185] In this embodiment of the disclosure, for different decisions, corresponding decision information features can be pre-set or learned. Therefore, the decision information obtained by decoding the first multimodal features can yield the corresponding decision information features. The decision information can be lateral control decisions and / or longitudinal control decisions, such as going straight, turning, accelerating, decelerating, etc.

[0186] In some alternative implementations, see [link to relevant documentation]. Figure 13 In the above 1204, the specific process by which the processor 21 determines the decision features based on multimodal features and decision information may include:

[0187] The decision information is encoded to obtain the decision information features (12041);

[0188] The multimodal features and decision information features are superimposed to obtain the decision features (12042).

[0189] In this process, decision information can be encoded to obtain decision information features, and then the decision information features and multimodal features can be superimposed to obtain the decision features.

[0190] In some implementations, decision information features and multimodal features can be added together, then fused through an MLP to output decision features; in other implementations, decision information features and multimodal features can be fused to output the fused decision features.

[0191] In the feature fusion of decision information features and multimodal features, the decision information features and multimodal features can be concatenated first, and then the concatenated features can be fused into a single decision feature using a Multilayer Perceptron (MLP); or, the decision information features and multimodal features can be added first, and then the concatenated features can be fused into a single decision feature using a Multilayer Perceptron (MLP).

[0192] Based on the embodiments of this disclosure, when vehicle trajectory prediction is required, environmental perception data, vehicle state data, and interaction data of the vehicle can be acquired first. Based on the vehicle state data and interaction data, text commands are determined. Then, based on the first sub-model (visual language model) in the trajectory planning model, the environmental perception data and text commands are processed to obtain multimodal features. Next, based on the multimodal features and decision information, decision features are determined. Finally, based on the second sub-model (end-to-end model) in the trajectory planning model, the decision features, environmental perception data, and vehicle state data are processed to obtain the vehicle's driving trajectory within a future time period. In this embodiment, the multimodal features output by the visual language model can provide fine-grained semantic information, while the decision information features corresponding to the decision category provide decision category information. By superimposing and fusing the multimodal features and decision information features and inputting them into the end-to-end model, both semantic information and decision category information can be provided to the end-to-end model simultaneously. This helps improve the accuracy of trajectory planning by the end-to-end model, provides accurate driving control basis for intelligent driving vehicles, and enhances the driving safety of intelligent driving vehicles.

[0193] Exemplary methods

[0194] The trajectory planning model generation method of this disclosure can be implemented by the trajectory planning model generation apparatus provided in any embodiment of this disclosure. The trajectory planning model generation method of this embodiment may include the following steps:

[0195] Obtain a training sample set, which includes at least one frame of input data and at least one frame of label data. The input data includes environmental perception data, agent state data, and interaction data. The label data includes trajectory labels and decision labels.

[0196] Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model.

[0197] In some optional examples, the initial trajectory planning model is trained in multiple stages based on the training sample set to obtain the trajectory planning model, including:

[0198] Based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models;

[0199] Based on the training sample set, the third sub-model and the fourth sub-model are jointly trained to obtain the first trajectory planning model;

[0200] Based on the training sample set, the first trajectory planning model is trained to obtain the trajectory planning model.

[0201] In some optional examples, based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models, including:

[0202] Based on agent state data and interaction data, determine text instructions;

[0203] Based on environmental perception data, text commands, and decision labels, the first sub-model in the initial trajectory planning model is trained to obtain the third sub-model;

[0204] Based on environmental perception data, agent state data, and trajectory labels, the second sub-model in the initial trajectory planning model is trained to obtain the fourth sub-model.

[0205] In some optional examples, the third sub-model and the fourth sub-model are connected through an initial connection layer;

[0206] Based on the training sample set, the third and fourth sub-models are jointly trained to obtain the first trajectory planning model, which includes:

[0207] Based on the training sample set, the initial trajectory planning model is trained to update the model parameters in the initial trajectory planning model except for the third sub-model, thus obtaining the second trajectory planning model.

[0208] The second trajectory planning model is trained based on the training sample set to obtain the first trajectory planning model.

[0209] In some optional examples, the first trajectory planning model is trained based on the training sample set to obtain a trajectory planning model, including:

[0210] Based on the first trajectory planning model, the first frame of input data in at least one frame of input data is iteratively processed to obtain the first predicted trajectory and at least one first prediction decision.

[0211] Based on the first prediction decision and the first prediction trajectory, the first trajectory planning model is trained to obtain the trajectory planning model.

[0212] In some optional examples, based on the training sample set, the initial trajectory planning model is trained to update the model parameters in the initial trajectory planning model excluding the third sub-model, resulting in a second trajectory planning model, including:

[0213] Based on the third sub-model in the initial trajectory planning model, the environmental perception data and text commands in the training samples are processed to obtain multimodal features and a second prediction decision.

[0214] Based on the initial connection layer in the initial trajectory planning model, multimodal features and second prediction decisions are processed to obtain decision features;

[0215] Based on the fourth sub-model in the initial trajectory planning model, the decision features and environmental perception data and agent state data in the training samples are processed to obtain the second predicted trajectory.

[0216] Based on the second predicted trajectory and trajectory label, the initial connection layer and the fourth sub-model in the initial trajectory planning model are trained to obtain the second trajectory planning model.

[0217] In some optional examples, the second trajectory planning model is trained based on the training sample set to obtain the first trajectory planning model, including:

[0218] Based on the second trajectory planning model, the training samples are processed to obtain the third predicted trajectory and the third predicted decision.

[0219] Based on kinematic constraints, the matching degree between the third predicted trajectory and the third predicted decision is determined;

[0220] In response to the mismatch between the matching degree characterization of the third predicted trajectory and the third predicted decision, the second trajectory planning model is trained based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model.

[0221] In some optional examples, the second trajectory planning model is trained based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model, including:

[0222] Based on the third prediction decision and the decision label, determine the value of the first loss function;

[0223] The value of the second loss function is determined based on the third predicted trajectory and trajectory label;

[0224] The first loss function value and the second loss function value are weighted and calculated to obtain the third loss function value;

[0225] The second trajectory planning model is trained based on the third loss function value to obtain the first trajectory planning model.

[0226] In some optional examples, based on the first trajectory planning model, the first frame of input data in at least one frame of input data is iteratively processed to obtain a first predicted trajectory of a preset duration and at least one first prediction decision, including:

[0227] Based on the first trajectory planning model, the first frame input data is processed to obtain the fourth predicted trajectory and the first prediction decision corresponding to the first frame input data;

[0228] Determine the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory;

[0229] Based on the 3D reconstruction model, the pose and time information corresponding to the first frame data are processed to render and generate a virtual scene image;

[0230] Based on the first trajectory planning model, the agent state data corresponding to the virtual scene image and the first frame data are processed to obtain the fourth predicted trajectory and the first predicted decision corresponding to the first frame data.

[0231] Repeatedly execute the operation of determining the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory to obtain the first predicted trajectory and at least one first predicted decision.

[0232] In some optional examples, a first trajectory planning model is trained based on a first prediction decision and a first predicted trajectory to obtain a trajectory planning model, including:

[0233] Based on the first predicted trajectory and the real environment data corresponding to at least one frame of input data, the collision risk value is determined.

[0234] In response to the collision risk value indicating the existence of collision risk, the first predicted trajectory and the first predicted decision are optimized to obtain the optimized predicted trajectory and the optimized predicted decision.

[0235] Based on the first predicted trajectory, the first predicted decision, the optimized predicted trajectory, and the optimized predicted decision, the first trajectory planning model is trained to obtain the trajectory planning model.

[0236] In some optional examples, the first trajectory planning model is trained based on the first predicted trajectory, the first predicted decision, the optimized predicted trajectory, and the optimized predicted decision to obtain a trajectory planning model, including:

[0237] Based on the first prediction decision and the optimized prediction decision, the value of the fourth loss function is determined;

[0238] Based on the first predicted trajectory and the optimized predicted trajectory, determine the value of the fifth loss function;

[0239] The sixth loss function value is obtained by weighting the fourth and fifth loss function values.

[0240] Based on the sixth loss function value, the first trajectory planning model is trained to obtain the trajectory planning model.

[0241] The method of this disclosure embodiment is implemented by the trajectory planning model generation device of any embodiment of this disclosure. The two are consistent in specific implementation and can be referenced and cited by each other. The specific implementation of the steps in the method of this disclosure embodiment can be implemented with reference to the above-mentioned device, and will not be repeated here.

[0242] The beneficial technical effects corresponding to the exemplary embodiments of this method can be found in the corresponding beneficial technical effects of the exemplary device section above, and will not be repeated here.

[0243] The vehicle trajectory planning method of this disclosure can be implemented by the vehicle trajectory planning device provided in any embodiment of this disclosure. The vehicle trajectory planning method of this embodiment may include the following steps:

[0244] Acquire vehicle environmental perception data, vehicle status data, and interaction data;

[0245] Based on vehicle status data and interaction data, determine the text instructions;

[0246] Based on the first sub-model in the trajectory planning model, the environmental perception data and text commands are processed to obtain the first multimodal features;

[0247] Based on the first multimodal features and decision information, the decision features are determined;

[0248] Based on the second sub-model in the trajectory planning model, the decision features, environmental perception data and vehicle state data are processed to obtain the vehicle's driving trajectory in the future time period.

[0249] In some optional examples, the first multimodal features are decoded to obtain decision information.

[0250] In some optional examples, the specific process of determining decision features based on multimodal features and decision information may include:

[0251] Decision information is encoded to obtain decision information characteristics;

[0252] The decision features are obtained by overlaying the multimodal features and decision information features.

[0253] In the methods of the embodiments of this disclosure, the various optional embodiments, optional implementation methods and optional examples disclosed in the above exemplary method section can be flexibly selected and combined as needed to achieve the corresponding functions and effects. This disclosure does not list them all.

[0254] The method of this disclosure embodiment is implemented by the vehicle trajectory planning device of any embodiment of this disclosure. The two are consistent in specific implementation and can be referenced and cited by each other. The specific implementation of the steps in the method of this disclosure embodiment can be implemented with reference to the above-mentioned device, and will not be repeated here.

[0255] The beneficial technical effects corresponding to the exemplary embodiments of this method can be found in the corresponding beneficial technical effects of the exemplary device section above, and will not be repeated here.

[0256] Exemplary electronic devices, computer program products, and computer-readable storage media

[0257] In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a structural diagram of an electronic device, including at least one processor and a memory.

[0258] A processor can be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and can control other components in an electronic device to perform desired functions.

[0259] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute one or more computer program instructions to implement the task-aware methods and / or other desired functions of the various embodiments of this disclosure described above.

[0260] In one example, the electronic device may also include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0261] The input device may also include, for example, a keyboard, a mouse, etc.

[0262] The output device can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0263] In addition, depending on the specific application, electronic devices may include any other suitable components.

[0264] Embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform steps in the vehicle trajectory planning methods of various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0265] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0266] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the vehicle trajectory planning methods of various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0267] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0268] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0269] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, this disclosure is also intended to include such modifications and variations if they fall within the scope of the claims of this disclosure and their equivalents.

Claims

1. An apparatus for generating a trajectory planning model, comprising a processor, the processor being configured to: Obtain a training sample set, wherein the training sample includes at least one frame of input data and at least one frame of label data, wherein the input data includes environmental perception data, agent state data, and interaction data, and the label data includes trajectory labels and decision labels; Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model.

2. The apparatus according to claim 1, wherein, The step of training the initial trajectory planning model in multiple stages based on the training sample set to obtain the trajectory planning model includes: Based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models. Based on the training sample set, the third sub-model and the fourth sub-model are jointly trained to obtain the first trajectory planning model; Based on the training sample set, the first trajectory planning model is trained to obtain the trajectory planning model.

3. The apparatus according to claim 2, wherein, Based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models, including: Based on the agent's state data and the interaction data, determine the text command; Based on the environmental perception data, the text instructions, and the decision labels, the first sub-model in the initial trajectory planning model is trained to obtain the third sub-model. Based on the environmental perception data, the agent state data, and the trajectory labels, the second sub-model in the initial trajectory planning model is trained to obtain the fourth sub-model.

4. The apparatus according to any one of claims 2-3, wherein, The third sub-model and the fourth sub-model are connected through an initial connection layer; The first trajectory planning model is obtained by jointly training the third sub-model and the fourth sub-model based on the training sample set, including: Based on the training sample set, the initial trajectory planning model is trained to update the model parameters in the initial trajectory planning model except for the third sub-model, to obtain the second trajectory planning model; Based on the training sample set, the second trajectory planning model is trained to obtain the first trajectory planning model.

5. The apparatus according to any one of claims 2-4, wherein, The step of training the first trajectory planning model based on the training sample set to obtain the trajectory planning model includes: Based on the first trajectory planning model, the first frame of input data in the at least one frame of input data is iteratively processed to obtain a first predicted trajectory and at least one first prediction decision. Based on the first prediction decision and the first prediction trajectory, the first trajectory planning model is trained to obtain the trajectory planning model.

6. The apparatus according to claim 4, wherein, The step of training the initial trajectory planning model based on the training sample set to update the model parameters in the initial trajectory planning model excluding the third sub-model, to obtain the second trajectory planning model, includes: Based on the third sub-model in the initial trajectory planning model, the environmental perception data and text instructions in the training samples are processed to obtain multimodal features and a second prediction decision; Based on the initial connection layer in the initial trajectory planning model, the multimodal features and the second prediction decision are processed to obtain decision features; Based on the fourth sub-model in the initial trajectory planning model, the decision features and the environmental perception data and agent state data in the training samples are processed to obtain the second predicted trajectory; Based on the second predicted trajectory and the trajectory label, the initial connection layer and the fourth sub-model in the initial trajectory planning model are trained to obtain the second trajectory planning model.

7. The apparatus according to claim 4, wherein, The step of training the second trajectory planning model based on the training sample set to obtain the first trajectory planning model includes: Based on the second trajectory planning model, the training samples are processed to obtain the third predicted trajectory and the third predicted decision; Based on kinematic constraints, the matching degree between the third predicted trajectory and the third predicted decision is determined; In response to the matching degree indicating a mismatch between the third predicted trajectory and the third predicted decision, the second trajectory planning model is trained based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model.

8. The apparatus according to claim 7, wherein, The process of training the second trajectory planning model based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model includes: Based on the third prediction decision and the decision label, the value of the first loss function is determined; Based on the third predicted trajectory and trajectory label, determine the value of the second loss function; The first loss function value and the second loss function value are weighted and calculated to obtain the third loss function value; Based on the third loss function value, the second trajectory planning model is trained to obtain the first trajectory planning model.

9. The apparatus according to claim 5, wherein, The step of iteratively processing the first frame of input data in the at least one frame of input data based on the first trajectory planning model to obtain a first predicted trajectory of a preset duration and at least one first prediction decision includes: Based on the first trajectory planning model, the first frame input data is processed to obtain the fourth predicted trajectory and the first prediction decision corresponding to the first frame input data. Determine the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory; Based on the 3D reconstruction model, the pose information and time information corresponding to the first frame data are processed to render and generate a virtual scene image; Based on the first trajectory planning model, the virtual scene image and the agent state data corresponding to the first frame data are processed to obtain the fourth predicted trajectory and the first predicted decision corresponding to the first frame data. Repeat the operation of determining the pose information, time information, and agent state data corresponding to the first frame data in the fourth predicted trajectory to obtain the first predicted trajectory and the at least one first prediction decision.

10. The apparatus according to claim 5, wherein, The step of training the first trajectory planning model based on the first prediction decision and the first prediction trajectory to obtain the trajectory planning model includes: Based on the real environment data corresponding to the first predicted trajectory and the at least one frame of input data, a collision risk value is determined. In response to the collision risk value indicating the existence of a collision risk, the first predicted trajectory and the first predicted decision are optimized to obtain an optimized predicted trajectory and an optimized predicted decision. Based on the first predicted trajectory, the first predicted decision, the optimized predicted trajectory, and the optimized predicted decision, the first trajectory planning model is trained to obtain the trajectory planning model.

11. The apparatus according to claim 10, wherein, The step of training the first trajectory planning model based on the first predicted trajectory, the first prediction decision, the optimized predicted trajectory, and the optimized prediction decision to obtain the trajectory planning model includes: Based on the first prediction decision and the optimized prediction decision, the value of the fourth loss function is determined; Based on the first predicted trajectory and the optimized predicted trajectory, determine the value of the fifth loss function; The sixth loss function value is obtained by weighting the fourth loss function value and the fifth loss function value. Based on the sixth loss function value, the first trajectory planning model is trained to obtain the trajectory planning model.

12. A vehicle trajectory planning device, comprising a processor, the processor being configured to: Acquire vehicle environmental perception data, vehicle status data, and interaction data; Based on the vehicle status data and the interaction data, determine the text command; Based on the first sub-model in the trajectory planning model, the environmental perception data and the text instructions are processed to obtain multimodal features; Based on the aforementioned multimodal features and decision information, decision features are determined; Based on the second sub-model in the trajectory planning model, the decision features, the environmental perception data, and the vehicle state data are processed to obtain the vehicle's driving trajectory in the future time period.

13. The apparatus of claim 12, wherein the processor is further configured to: The multimodal features are decoded to obtain decision information; Based on the decision information, the features of the decision information are obtained.

14. The apparatus according to any one of claims 12-13, wherein, The step of determining decision features based on the multimodal features and decision information includes: The decision information is encoded to obtain decision information features; The decision features are obtained by superimposing the multimodal features and the decision information features.

15. The apparatus according to any one of claims 12-14, wherein, The trajectory planning model is obtained by training the device described in any one of claims 1-11.

16. A method for generating a trajectory planning model, comprising: Obtain a training sample set, wherein the training sample includes at least one frame of input data and at least one frame of label data, wherein the input data includes environmental perception data, agent state data, and interaction data, and the label data includes trajectory labels and decision labels; Based on the training sample set, the initial trajectory planning model is trained in multiple stages to obtain the trajectory planning model.

17. The method according to claim 16, wherein, The step of training the initial trajectory planning model in multiple stages based on the training sample set to obtain the trajectory planning model includes: Based on the training sample set, the first and second sub-models in the initial trajectory planning model are trained respectively to obtain the third and fourth sub-models. Based on the training sample set, the third sub-model and the fourth sub-model are jointly trained to obtain the first trajectory planning model; Based on the training sample set, the first trajectory planning model is trained to obtain the trajectory planning model.

18. The method according to claim 17, wherein, The step of training the first trajectory planning model based on the training sample set to obtain the trajectory planning model includes: Based on the first trajectory planning model, the first frame of input data in the at least one frame of input data is iteratively processed to obtain a first predicted trajectory and at least one first prediction decision. Based on the first prediction decision and the first prediction trajectory, the first trajectory planning model is trained to obtain the trajectory planning model.

19. The method of claim 17, wherein, The third sub-model and the fourth sub-model are connected through an initial connection layer; The first trajectory planning model is obtained by jointly training the third sub-model and the fourth sub-model based on the training sample set, including: Based on the training sample set, the initial trajectory planning model is trained to update the model parameters in the initial trajectory planning model except for the third sub-model, to obtain the second trajectory planning model; Based on the training sample set, the second trajectory planning model is trained to obtain the first trajectory planning model.

20. The method according to claim 19, wherein, The step of training the second trajectory planning model based on the training sample set to obtain the first trajectory planning model includes: Based on the second trajectory planning model, the training samples are processed to obtain the third predicted trajectory and the third predicted decision; Based on kinematic constraints, the matching degree between the third predicted trajectory and the third predicted decision is determined; In response to the matching degree indicating a mismatch between the third predicted trajectory and the third predicted decision, the second trajectory planning model is trained based on the third predicted trajectory, the third predicted decision, the trajectory label, and the decision label to obtain the first trajectory planning model.

21. A method for planning vehicle trajectories, comprising: Acquire vehicle environmental perception data, vehicle status data, and interaction data; Based on the vehicle status data and the interaction data, determine the text command; Based on the first sub-model in the trajectory planning model, the environmental perception data and the text instructions are processed to obtain multimodal features; Based on the aforementioned multimodal features and decision information, decision features are determined; Based on the second sub-model in the trajectory planning model, the decision features, the environmental perception data, and the vehicle state data are processed to obtain the vehicle's driving trajectory in the future time period.

22. A computer-readable storage medium storing computer program instructions that, when executed, implement the method described in any one of claims 16-21.

23. A computer program product comprising computer program instructions, which, when executed by a processor, implement the method described in any one of claims 16-21.

24. An electronic device, the electronic device comprising: Memory, used to store computer program products; A processor is configured to execute a computer program product stored in the memory, wherein when the computer program product is executed, it implements the method for generating a trajectory planning model as described in any one of claims 16-20, or the method for planning a vehicle trajectory as described in claim 21.