Training method, trajectory planning method and electronic equipment

By optimizing the trajectory planning model based on visual data prediction and a multi-stage reasoning process, the shortcomings of existing models in trajectory planning in complex traffic scenarios are addressed, more accurate and efficient trajectory planning is achieved, and the performance and reliability of the autonomous driving system are improved.

CN120652866APending Publication Date: 2025-09-16EACON TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510660244.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing trajectory planning models are difficult to adapt to complex and changing traffic conditions in autonomous driving, and are unable to achieve efficient and accurate end-to-end trajectory planning, especially in terms of context modeling, complex task decomposition and causal reasoning.

Method used

By predicting the visual data of the next scene based on the actual visual data of the current scene, the first adjustment target parameters are determined, and a multi-stage reasoning process is performed by combining the actual visual data, command data and predicted visual data to generate the target vehicle trajectory planning. The mean square error and cross entropy loss functions are used to adjust the trajectory planning model and optimize its parameters to adapt to complex traffic scenarios.

Benefits of technology

It improves the trajectory planning model's perception and prediction accuracy of environmental changes, enhances its reasoning and decision-making capabilities in complex traffic scenarios, improves the efficiency and accuracy of end-to-end trajectory planning, and improves the overall performance and reliability of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120652866A_ABST
    Figure CN120652866A_ABST
Patent Text Reader

Abstract

The invention provides a training method, a trajectory planning method and electronic equipment, and relates to the technical field of automatic driving. The training method comprises the following steps: determining predicted visual data of a next scene based on actual visual data of a current scene; determining a first adjustment target parameter according to the predicted visual data of the next scene and the actual visual data of the next scene; determining a multi-stage reasoning process based on the actual visual data of the current scene, the instruction data and the predicted visual data of the next scene, and generating a target vehicle trajectory plan; determining a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning and the multi-stage reasoning process answer and the target vehicle trajectory planning answer of the next scene; and adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter. According to the method, the trajectory planning model can more accurately adapt to complex and changeable traffic conditions, so that the accuracy of end-to-end trajectory planning is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and specifically to a training method, a trajectory planning method, and an electronic device. Background Art

[0002] In applications in the field of autonomous driving, it is necessary to continuously perceive environmental changes and make inferences and decisions based on the context, which places higher demands on the model's cognitive ability, logical reasoning ability, and behavior planning ability.

[0003] However, existing trajectory planning models are still insufficient in terms of context modeling, complex task decomposition, and causal reasoning in autonomous driving tasks. They are unable to adapt to complex and changing traffic conditions and cannot achieve end-to-end efficient and accurate trajectory planning. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a training method, a trajectory planning method, and an electronic device.

[0005] In a first aspect, an embodiment of the present application provides a training method for training a trajectory planning model, the method comprising: determining predicted visual data for a next scene based on actual visual data of a current scene; determining a first adjustment target parameter based on the predicted visual data for the next scene and the actual visual data for the next scene; determining a multi-stage reasoning process based on the actual visual data of the current scene, instruction data, and predicted visual data for the next scene, and generating a target vehicle trajectory planning based on the multi-stage reasoning process and the predicted visual data for the next scene; determining a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning, the answer to the multi-stage reasoning process for the next scene, and the answer to the target vehicle trajectory planning; and adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter.

[0006] In combination with the first aspect, in certain implementations of the first aspect, a first adjustment target parameter is determined based on the predicted visual data of the next scene and the actual visual data of the next scene, including: using a mean square error loss function to determine feature difference data between the predicted visual data of the next scene and the actual visual data of the next scene; and determining the feature difference data as the first adjustment target parameter.

[0007] In combination with the first aspect, in certain implementations of the first aspect, the second adjustment target parameter is determined based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer of the next scene, including: determining the cross entropy loss by using the cross entropy loss function based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer of the next scene; determining the cross entropy loss as the second adjustment target parameter.

[0008] In combination with the first aspect, in certain implementations of the first aspect, the trajectory planning model includes a scene prediction module and a trajectory planning module, and adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter includes: adjusting the scene prediction module and the trajectory planning module to minimize the sum of the first adjustment target parameter and the second adjustment target parameter.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the actual visual data of the current scene includes an actual front view image, the predicted visual data of the next scene includes a predicted front view image, and determining the predicted visual data of the next scene based on the actual visual data of the current scene includes: segmenting the actual front view image of the current scene into multiple grid images; and determining the predicted front view image of the next scene based on the multiple grid images and the actual front view image of the current scene.

[0010] In combination with the first aspect, in certain implementations of the first aspect, the predicted visual data of the next scene is determined based on the actual visual data of the current scene, including: extracting features from the actual visual data of the current scene to obtain visual features; determining the driving information and navigation instructions of the target vehicle, and determining the contextual features corresponding to the driving information and navigation instructions; fusing the visual features and the contextual features to obtain a first fused feature; and determining the predicted visual data of the next scene based on the first fused feature.

[0011] In combination with the first aspect, in certain implementations of the first aspect, a multi-stage reasoning process is determined based on the actual visual data, instruction data of the current scene, and the predicted visual data of the next scene, including: fusing the actual visual data and instruction data of the current scene to generate a second fusion feature, the instruction data including navigation instruction data and trajectory prediction instruction data; based on the second fusion feature and the predicted visual data of the next scene, a multi-stage reasoning process is determined, the multi-stage reasoning process including at least one of scene understanding, traffic sign recognition, target obstacle recognition, and driving risk assessment.

[0012] In combination with the first aspect, in certain implementations of the first aspect, the instruction data includes instruction data in text form, and the actual visual data of the current scene and the instruction data are fused to generate a second fusion feature, including: determining multiple text units corresponding to the instruction data, and visual placeholder symbols contained in the multiple text units; replacing the visual placeholder symbols with the actual visual data of the current scene associated with the semantic features of the target text unit to generate a second fusion feature, where the target text unit represents the text unit associated with the position of the visual placeholder symbol.

[0013] In combination with the first aspect, in certain implementations of the first aspect, the multi-stage reasoning process includes meta-action decision-making, and generating target vehicle trajectory planning based on the multi-stage reasoning process and predicted visual data of the next scene, including: generating target vehicle trajectory planning based on meta-action decision-making and predicted visual data of the next scene.

[0014] In the second aspect, an embodiment of the present application provides a trajectory planning method, including: obtaining visual data and command data of the current scene collected by the target vehicle; using a trajectory planning model to process the visual data and command data of the current scene to obtain a multi-stage reasoning process and target vehicle trajectory planning, and the trajectory planning model is trained based on the method described in the first aspect.

[0015] In a third aspect, an embodiment of the present application provides a training device for training a trajectory planning model, the device comprising: a first determination module for determining the predicted visual data of the next scene based on the actual visual data of the current scene; a second determination module for determining a first adjustment target parameter based on the predicted visual data of the next scene and the actual visual data of the next scene; a third determination module for determining a multi-stage reasoning process based on the actual visual data of the current scene, instruction data and the predicted visual data of the next scene, and generating a target vehicle trajectory planning based on the multi-stage reasoning process and the predicted visual data of the next scene; a fourth determination module for determining a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning and the answer to the multi-stage reasoning process and the target vehicle trajectory planning answer for the next scene; and an adjustment module for adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter.

[0016] In a fourth aspect, an embodiment of the present application provides a trajectory planning device, including: an acquisition module for acquiring visual data and instruction data of the current scene collected by the target vehicle; a processing module for using a trajectory planning model to process the visual data and instruction data of the current scene to obtain a multi-stage reasoning process and target vehicle trajectory planning, and the trajectory planning model is trained based on the method described in the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program for executing the methods described in the first and second aspects.

[0018] In a sixth aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; and the processor is configured to execute the methods described in the first and second aspects.

[0019] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed on an electronic device, the electronic device implements the method described in the first and second aspects.

[0020] In this application, by predicting the visual data of the next scene based on the actual visual data of the current scene and determining the first adjustment target parameter accordingly, the trajectory planning model's perception and prediction accuracy of environmental changes are effectively improved, thereby achieving more comprehensive scene understanding and context-based forward-looking predictions. Furthermore, by combining actual visual data, command data, and predicted visual data, a multi-stage reasoning process is determined and a target vehicle trajectory plan is generated, further enhancing the trajectory planning model's reasoning and decision-making capabilities in complex traffic scenarios and the model's interpretability. Furthermore, by generating the target vehicle trajectory plan based on the multi-stage reasoning process and the predicted visual data of the next scene, the model is forward-looking in its decision-making, breaking through the limitation of traditional models that rely solely on historical data. Furthermore, by determining the second adjustment target parameter based on the answers to the multi-stage reasoning process and the trajectory plan answer, it optimizes the trajectory planning model's reasoning and planning decisions. Ultimately, by combining the first and second adjustment target parameters to adjust the trajectory planning model, the trajectory planning model can more accurately adapt to complex and changing traffic conditions, significantly improving the efficiency and accuracy of end-to-end trajectory planning and enhancing the overall performance and reliability of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 The figure is a flow chart of a training method provided in one embodiment of the present application.

[0023] Figure 2 The figure is a flow chart of a trajectory planning method provided in one embodiment of the present application.

[0024] Figure 3 Shown is a schematic structural diagram of a training device provided in one embodiment of the present application.

[0025] Figure 4 Shown is a structural schematic diagram of a trajectory planning device provided in one embodiment of the present application.

[0026] Figure 5 Shown is a structural schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] Figure 1 FIG. 1 is a flow chart of a training method according to an embodiment of the present application. Figure 1 As shown, the method includes the following steps.

[0029] Step S110 : determining predicted visual data of the next scene based on the actual visual data of the current scene.

[0030] The actual visual data of the current scene refers to the real road environment image information collected by the vehicle's camera, lidar and other sensors at the current moment during the target vehicle's driving process, including the shape of the road, the positions of surrounding vehicles and pedestrians, traffic signs and signals, and other visual content related to the driving scene. Optionally, the actual visual data of the current scene includes multi-view RGB images, denoted as Where N represents the number of viewing angles, W and H are the width and height of the RGB image respectively.

[0031] Accordingly, the predicted visual data of the next scene is the visual observation content of the target vehicle in the scene at the next moment, which is inferred based on the actual visual data of the current scene, so as to estimate the future road environment image situation. Optionally, the predicted visual data of the next scene and the actual visual data of the current scene are represented in the same form.

[0032] In one implementation, actual visual data of the current scene is collected, and then feature extraction and encoding are performed on this data to capture key visual information. Simultaneously, the predicted visual data for the next scene is derived by combining the target vehicle's kinematic model with dynamic parameters such as its current speed, acceleration, and steering angle.

[0033] Step S120 : determining a first adjustment target parameter according to the predicted visual data of the next scene and the actual visual data of the next scene.

[0034] The first adjustment target parameter refers to a target parameter determined based on the difference between the predicted visual data and the actual visual data and used to adjust the relevant parameters of the trajectory planning model (such as weights, biases, etc.).

[0035] In one implementation, the actual visual data of the next scene is used as the target frame. Optionally, after AnyRes grid segmentation, SigLIP encoder and MLP projection processing, a potential visual representation of the actual visual data of the next scene is obtained. This visual representation is used as a self-supervisory signal and compared with the visual representation corresponding to the predicted visual data to calculate the degree of difference between the two. This difference reflects the lack of accuracy of the trajectory planning model's prediction. Then, based on this degree of difference, the first adjustment target parameter is determined so that the trajectory planning model can better fit the actual visual data.

[0036] Step S130 , based on the actual visual data of the current scene, the instruction data, and the predicted visual data of the next scene, a multi-stage reasoning process is determined, and a target vehicle trajectory plan is generated according to the multi-stage reasoning process and the predicted visual data of the next scene.

[0037] Command data refers to navigation instructions or trajectory prediction instructions related to the current driving task, such as "turn left" or "go straight through the intersection," providing guidance for the target vehicle's driving direction and destination. The multi-stage reasoning process breaks down the target vehicle's decision-making process into multiple, sequential stages, each completing a specific task, gradually progressing from environmental perception to final decision-making.

[0038] In one implementation, the system uses the actual visual data of the current scene as a foundation, combined with command data, to determine the target vehicle's driving destination and direction. It also references the predicted visual data of the next scene to anticipate future environmental changes. Based on this information, a multi-stage reasoning process is initiated. Finally, based on this multi-stage reasoning process and the predicted visual data of the next scene, a path planning algorithm is used to generate a trajectory plan for the target vehicle, ensuring that the target vehicle can safely and efficiently reach its target location.

[0039] Step S140 , determining a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer for the next scenario.

[0040] The answer to the multi-stage reasoning process for the next scenario refers to the verified, correct results of the multi-stage reasoning process in the next scenario, which can be used as a reference standard for evaluating and adjusting the current multi-stage reasoning process. The answer to the target vehicle trajectory plan refers to the target vehicle's actual driving trajectory or the verified optimal trajectory in the next scenario, which is used to evaluate and adjust the rationality of the currently generated target vehicle trajectory plan.

[0041] In one implementation, a second adjustment target parameter is determined based on the difference between the multi-stage reasoning process and the answer to the multi-stage reasoning process, and the difference between the target vehicle trajectory plan and the answer to the target vehicle trajectory plan. Similarly, the second adjustment target parameter is a target parameter used to optimize parameters related to the trajectory planning model, with the goal of making the reasoning and planning results of the trajectory planning model closer to reality.

[0042] Step S150 : adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter.

[0043] The first adjustment target parameter and the second adjustment target parameter point out the current deviations and improvement directions of the trajectory planning model from the two dimensions of visual prediction and reasoning planning, respectively. Optionally, in this embodiment, the two adjustment target parameters are combined, and the weights, biases and other parameters in the trajectory planning model are comprehensively adjusted through an optimization algorithm (such as gradient descent, etc.). Specifically, according to the first adjustment target parameter, the parts of the trajectory planning model related to visual feature extraction and time series prediction are optimized to improve the prediction accuracy of the trajectory planning model for the visual data of the next scene. At the same time, according to the second adjustment target parameter, the parameters related to multi-stage reasoning and decision generation in the trajectory planning model are adjusted to enhance the reasoning ability of the trajectory planning model in complex scenarios and the rationality of trajectory planning.

[0044] This application addresses the shortcomings of existing trajectory planning models in the field of autonomous driving in terms of context modeling, complex task decomposition, and causal reasoning, and proposes a novel training method. Specifically, by predicting the visual data of the next scene based on the actual visual data of the current scene and determining the first adjustment target parameter accordingly, the trajectory planning model's perception and prediction accuracy of environmental changes are effectively improved, thereby achieving more comprehensive scene understanding and context-based forward-looking predictions. At the same time, by combining actual visual data, instruction data, and predicted visual data, a multi-stage reasoning process is determined and the target vehicle trajectory plan is generated, further enhancing the trajectory planning model's reasoning and decision-making capabilities in complex traffic scenarios and the model's interpretability. Furthermore, by generating the target vehicle trajectory plan based on the multi-stage reasoning process and the predicted visual data of the next scene, the model is forward-looking when making decisions, breaking through the limitation of traditional models that rely solely on historical data. In addition, the second adjustment target parameter is determined based on the answers to the multi-stage reasoning process and the trajectory plan answer, achieving optimized adjustment of the trajectory planning model's reasoning and planning decisions. Ultimately, by combining the first and second adjustment target parameters to adjust the trajectory planning model, the trajectory planning model can more accurately adapt to complex and changing traffic conditions, thereby significantly improving the efficiency and accuracy of end-to-end trajectory planning and enhancing the overall performance and reliability of the autonomous driving system.

[0045] Figure 1The training method in the illustrated embodiment improves the performance of the trajectory planning model in autonomous driving tasks through multi-step collaborative optimization, enhancing its ability to cope with complex traffic scenarios. The following details how to determine the first and second adjustment target parameters, and how to adjust the trajectory planning model based on these parameters.

[0046] In some embodiments, a mean square error loss function is used to determine feature difference data between the predicted visual data of the next scene and the actual visual data of the next scene; and the feature difference data is determined as the first adjustment target parameter.

[0047] The mean squared error loss function is a function used to measure the difference between the predicted value and the true value. It quantifies the error by calculating the average of the squares of the differences between the predicted value and the true value.

[0048] In one implementation, first, the predicted visual data and actual visual data of the next scene are obtained. In order to quantify the difference between the two sets of data, the mean square error loss function is used to calculate the feature difference data between the two. Specifically, by solving the average value of the square of the difference between the predicted visual data and the actual visual data on the corresponding feature dimension, a scalar value, namely the feature difference data, is obtained. This data intuitively reflects the accuracy deviation of the trajectory planning model when predicting the visual data of the next scene. After determining the feature difference data, it is used as the first adjustment target parameter. Subsequently, the trajectory planning model can be optimized and adjusted in a targeted manner based on this parameter to reduce the prediction error and improve the trajectory planning model's ability to predict the visual data of the next scene.

[0049] In this embodiment, the mean squared error loss function is used to measure the difference between predicted and actual visual data, enabling the trajectory planning model to accurately identify deviations in the prediction process. Accordingly, this quantified feature difference data directly reflects the trajectory planning model's deficiencies in visual prediction, enabling targeted adjustment and optimization of relevant parameters during subsequent model training, thereby enhancing the trajectory planning model's performance in the critical aspect of perceiving environmental changes.

[0050] In some embodiments, based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer for the next scene, a cross-entropy loss function is used to determine the cross-entropy loss; and the cross-entropy loss is determined as the second adjustment target parameter.

[0051] The cross entropy loss function is a function used to measure the difference between the model's predicted results and the actual results. It is often used in classification problems to quantify the error by calculating the difference between the predicted probability distribution and the true probability distribution.

[0052] Optionally, the multi-stage reasoning process, the target vehicle trajectory planning and the corresponding answers are input into the cross-entropy loss function together. The cross-entropy loss function compares the difference between the reasoning process and trajectory planning predicted by the trajectory planning model (expressed in the form of probability distribution) and the actual answer (i.e., the true distribution), and calculates a scalar value reflecting this difference, namely the cross-entropy loss. Exemplarily, the answer to the multi-stage reasoning process includes the explicit supervision corresponding to each of the multi-stage reasoning processes, that is, each stage of the reasoning process has a standard answer to supervise the learning of the trajectory planning model, and the explicit supervision can be used as the answer to the multi-stage reasoning process when calculating the cross-entropy loss.

[0053] In other embodiments, when calculating the cross-entropy loss, the original input of the trajectory planning model (i.e., the actual visual data of the current scene), the predicted visual data of the next scene, and their respective answers can also be input into the cross-entropy loss function to calculate the cross-entropy loss, and the cross-entropy loss is used together with the previously calculated cross-entropy loss as the cross-entropy loss of the trajectory planning model to determine the second adjustment target parameter.

[0054] The cross-entropy loss quantifies the degree of deviation between the trajectory planning model's predictions and the actual results under the current parameters. The smaller the cross-entropy loss, the closer the trajectory planning model's predictions are to reality. Finally, the calculated cross-entropy loss is determined as the second adjustment target parameter, which is used to guide the adjustment of the trajectory planning model parameters to minimize the cross-entropy loss.

[0055] In this embodiment, the cross-entropy loss function is used to effectively evaluate the difference between the model prediction and the actual results, providing a basis for parameter adjustment, thereby improving the performance and accuracy of the model in multi-stage reasoning and trajectory planning tasks, enhancing the model's adaptability to complex traffic scenarios and decision-making rationality, and improving the accuracy and reliability of trajectory planning.

[0056] In some embodiments, the scene prediction module and the trajectory planning module are adjusted to minimize the sum of the first adjustment target parameter and the second adjustment target parameter.

[0057] That is, the optimization process of the trajectory planning model involves the coordinated adjustment of the scene prediction module and the trajectory planning module. Optionally, through the model optimization algorithm, the parameters of the two modules are iteratively updated to minimize their sum while considering the first adjustment target parameter and the second adjustment target parameter. For example, when there is a deviation between the predicted visual data and the actual visual data of the next scene (measured by the first adjustment target parameter) and the trajectory planning does not match the actual optimal path (measured by the second adjustment target parameter), the model will calculate the contribution of these two modules to the total loss separately through the back propagation algorithm. Then, based on these contributions, the scene prediction module can adjust its network weights to more accurately capture environmental changes; the trajectory planning module optimizes its decision logic to generate a more reasonable trajectory planning. This process continues until the sum of the first adjustment target parameter and the second adjustment target parameter is minimized, thereby improving the overall performance of the model and ensuring that the prediction and planning results are closer to actual needs.

[0058] Understandably, when training a trajectory planning model, it's desirable to excel in both visual prediction and reasoning planning. The first and second adjustment target parameters represent the optimization objectives for these two areas, respectively. By minimizing their sum, these two objectives can be collaboratively optimized, avoiding the need to optimize one objective while neglecting the other. This integrated optimization approach ensures optimal model performance in both perception and decision-making planning, thereby improving the model's adaptability and reliability in complex traffic environments and ensuring safe and efficient autonomous driving.

[0059] Below, in Figure 1 On the basis of the illustrated embodiment, how to determine the predicted visual data of the next scene based on the actual visual data of the current scene is further refined from the perspectives of image processing and feature fusion.

[0060] In some embodiments, the actual visual data of the current scene includes an actual front view image, and the predicted visual data of the next scene includes a predicted front view image. Specifically, the actual front view image of the current scene is divided into multiple grid images; based on the multiple grid images and the actual front view image of the current scene, the predicted front view image of the next scene is determined.

[0061] A grid image refers to an image unit obtained by dividing the actual front view image into multiple small blocks according to certain rules. Each grid image represents the visual information of a part of the original image.

[0062] Specifically, first, the actual front view image of the current scene is segmented and divided into multiple grid images. This decomposes the complex image information into smaller, more easily processed parts, enabling the model to more accurately capture and understand the local features and details in the image. It should be noted that the number of grid images can be set according to actual processing requirements, and this application does not limit the specific number of grid images. Optionally, in this embodiment, the actual visual image of the current scene is segmented into four grid images.

[0063] Next, the trajectory planning model comprehensively analyzes these grid images and the actual front view image of the current scene to determine the predicted front view image for the next scene. During this process, the model considers the changing trends of each grid image in the time series and their correlation with the overall image. For example, if a grid image contains an approaching vehicle, the model will predict the vehicle's position and status at future moments based on its current motion state and trajectory, and integrate these predictions into the corresponding position in the predicted front view image of the next scene. At the same time, the model also combines global information in the actual front view image, such as the direction of the road and the status of traffic signals, to ensure the overall consistency and rationality of the prediction results.

[0064] It is understandable that in autonomous driving tasks, the front view image contains the richest semantic information and can capture key clues related to trajectory planning. Therefore, in order to improve training efficiency and reduce redundant calculations, this embodiment divides the actual front view image into multiple grid images. On the one hand, it improves the processing efficiency of the model. On the other hand, it also allows the model to more carefully capture and understand the local features in the image, so that the prediction results can more accurately reflect the changes in various parts of the future scene. In addition, the actual front view image of the current scene can provide rich contextual information for the subsequent reasoning process. Therefore, based on the comprehensive analysis of the grid image and the overall image, the model can maintain a grasp of the overall scene and ensure the global consistency of the predicted visual image, thereby providing a more accurate visual basis for subsequent trajectory planning.

[0065] In some embodiments, feature extraction is performed on actual visual data of the current scene to obtain visual features; driving information and navigation instructions of the target vehicle are determined, as well as context features corresponding to the driving information and navigation instructions are determined; the visual features and context features are fused to obtain a first fused feature; and based on the first fused feature, predicted visual data for the next scene is determined.

[0066] Specifically, feature extraction is performed on the actual visual data of the current scene to obtain visual features, that is, the complex image information is converted into feature vectors that can be processed by the model, and the key visual information in the scene is retained. Next, the driving information (such as speed, acceleration, etc.) and navigation instructions (such as driving direction, target lane, etc.) of the target vehicle are determined, and the corresponding context features are determined based on this information to provide the model with background information on the status of the target vehicle and the driving task. Optionally, the context encoder (such as a two-layer MLP module) in the trajectory planning model is used to embed the driving information and navigation instructions into a context representation, that is, a context feature. Then, the visual features and the context features are fused to obtain a first fused feature to enhance the expressive power of the feature. Finally, the predicted visual data for the next scene is determined based on the first fused feature.

[0067] For example, assuming that the actual visual data of the current scene is an actual front view image, the actual front view image is recorded as Adopt AnyRes strategy, Segmented into multiple spatial images. In the trajectory planning model, each grid image is processed by the visual encoder SigLIP to obtain visual features Among them L v represents the number of visual tokens in each grid image, D v is the visual embedding dimension.

[0068] Optionally, if the trajectory planning model is a large language model, it is also necessary to align the visual features to the language space. For example, through a two-layer multi-layer perceptron projection module, Map to Among them D p Embedding the language dimension enables visual features and text information to interact and fuse in the same semantic space.

[0069] In this embodiment, feature extraction is performed on the actual visual data of the current scene to obtain visual features. This step can retain the key visual information in the current scene. Simultaneously, the target vehicle's driving information and navigation instructions are determined, and contextual features are determined based on this information, enabling the model to understand the current driving intention and vehicle status. These two features are fused to obtain a first fused feature, allowing the model to comprehensively consider the visual information of the current scene and the vehicle's driving status, enhancing the feature's expressive power. Finally, based on this fused feature, the predicted visual data for the next scene is determined. This prediction is based not only on the current visual information but also incorporates the vehicle's driving intention and status, thereby more accurately reflecting future scenarios.

[0070] Below, in Figure 1Building on the illustrated embodiment, this paper further details how to determine a multi-stage reasoning process based on the actual visual data of the current scene, the instruction data, and the predicted visual data of the next scene. Specifically, the actual visual data of the current scene and the instruction data are fused to generate a second fused feature; the multi-stage reasoning process is determined based on the second fused feature and the predicted visual data of the next scene.

[0071] Optionally, the instruction data includes navigation instruction data (such as driving direction, target location, etc.) and trajectory prediction instruction data, providing mission and target information for the driving of the target vehicle.

[0072] Exemplarily, the actual visual data of the current scene is first fused with the instruction data. For example, when driving on a city road, the target vehicle's camera, lidar and other sensors collect actual visual data such as road images and the positions of surrounding vehicles in real time. At the same time, the on-board navigation system and trajectory prediction module provide navigation instruction data (such as turning left) and trajectory prediction instruction data (such as driving along a certain lane). Then, based on these data, a second fusion feature is generated through feature extraction and fusion algorithms (such as multimodal fusion technology in deep learning). Next, based on the second fusion feature and the predicted visual data of the next scene, a multi-stage reasoning process is determined. Optionally, the multi-stage reasoning process includes at least one of scene understanding, traffic sign recognition, target obstacle recognition, and driving risk assessment.

[0073] It's understandable that the predicted visual data for the next scene can provide forward-looking information for reasoning, helping to identify potential risks and environmental changes in advance, thereby enabling early decision-making and planning. Furthermore, this forward-looking information makes the reasoning process more coherent and accurate, avoiding misjudgments caused by localized visual information. At the same time, predicted visual data supports the gradual deployment of multi-stage reasoning, providing temporal continuity for each reasoning step, helping the trajectory planning model to more comprehensively understand the direction of the scene's development. Furthermore, it can improve the trajectory planning model's adaptability and robustness to complex environments, supplement the deficiencies of current visual information, and maintain stable decision-making capabilities.

[0074] For example, the navigation instruction data in the second fused feature indicates that the target vehicle is about to enter an intersection, while the predicted visual data indicates that a pedestrian may be crossing the road ahead. Based on this, the trajectory planning model initiates target obstacle identification and driving risk assessment in a multi-stage reasoning process, predicting the pedestrian's movement trajectory and assessing potential collision risk. In this way, the trajectory planning model ensures that the reasoning process in complex traffic scenarios not only meets the current mission objectives but also adapts to future changes in the scene, generating reasonable and safe driving decisions.

[0075] In this embodiment, the second fusion feature integrates the actual visual data and command data of the current scene. This not only contains rich environmental perception information, but also incorporates task-oriented information such as navigation and trajectory prediction, providing a comprehensive description of the current state for the reasoning process. The predicted visual data of the next scene injects foresight into the reasoning process, enabling the trajectory planning model to make judgments and decisions in advance based on expected future visual information. This combination ensures the temporal and logical coherence of the reasoning process, allowing each stage of reasoning to be carried out based on more complete spatiotemporal information, thereby improving the depth and accuracy of understanding of complex traffic scenarios.

[0076] In the process of fusing the actual visual data of the current scene with the instruction data to generate the second fused feature, we further explore a special case: when the instruction data exists in textual form. Specifically, we determine the multiple text units corresponding to the instruction data and the visual placeholders contained in the multiple text units; then, we replace the visual placeholders with the actual visual data of the current scene associated with the semantic features of the target text unit to generate the second fused feature. The target text unit represents the text unit associated with the position of the visual placeholder.

[0077] The visual placeholder represents a special identifier used to indicate the location where the visual data needs to be inserted in the text-based instruction data, so as to clearly point out which parts of the text-based instruction data need to be combined with the actual visual data, thereby achieving the fusion of text and visual data.

[0078] Exemplarily, the visual placeholder is recorded as a special token. First, the special tokens in the instruction data are identified, and these special tokens are used to indicate the position of the visual placeholder symbol. For example, the instruction data is "Please turn left based on the front view (special token) and left view (special token) of the main vehicle", where the special tokens correspond to the actual visual data such as the front view and the left view respectively. Then, the instruction data is divided into multiple text units based on the visual placeholder, such as "Please turn left based on the front view of the owner", "left view", and "turn left". According to the position of each special token and the semantic features of the corresponding target text unit, the actual visual data of the corresponding current scene is embedded into the represented instruction data, replacing the visual placeholder symbol, and then, a second fusion feature containing visual information is generated. For example, the front view image token is embedded in the position of the special token after the front view, and the left view image token is embedded in the position of the special token after the left view, thereby realizing the organic fusion of text and visual data and providing richer input information for the multi-stage reasoning process.

[0079] In this embodiment, multiple text units in the instruction data and the visual placeholders they contain are first determined, and then the visual placeholders are replaced with actual visual data semantically associated with the target text unit to generate a second fused feature. This ensures that each visual placeholder in the text-based instruction data corresponds to accurate visual information in the current scene, so that the fused feature retains the semantics of the text instruction and integrates key visual elements, so that the model can more effectively execute the multi-stage reasoning process, thereby improving the autonomous driving system's execution capability and decision-making accuracy for complex instructions.

[0080] exist Figure 1 In the training method shown, the multi-stage reasoning process is the key step in generating the target vehicle trajectory plan. In particular, when the multi-stage reasoning process includes meta-action decisions, the target vehicle trajectory plan is generated based on the meta-action decisions and the predicted visual data of the next scene.

[0081] Specifically, meta-action decision-making refers to the process of decomposing complex driving tasks into a series of basic, executable driving action decision-making steps in a multi-stage reasoning process. These basic actions are called "meta-actions". They are the smallest operation units that constitute complex driving behaviors, such as acceleration, deceleration, lane changing, turning, and parking.

[0082] The predicted visual data of the next scene provides information about the environment that the target vehicle may encounter in the future, such as changes in the road, the positions and movement trends of other vehicles and pedestrians, etc. By fusing the meta-action decision and the predicted visual data of the next scene, the model can predict the possible position and surrounding environment state of the target vehicle after performing the selected meta-action. Based on this information, an optimal trajectory is generated as the target vehicle trajectory planning to guide the target vehicle in the current scene and prepare for the next scene. For example, if the meta-action decision is "change lanes to the left", a smooth and safe lane change trajectory will be calculated based on the current vehicle speed, the distance to the adjacent vehicle, and the predicted movement trends of other vehicles in the next scene.

[0083] In this embodiment, meta-action decision-making breaks down complex driving tasks into basic actions. Then, using predicted visual data to predict future environmental changes, the trajectory planning model plans the corresponding driving path based on these meta-action decisions. This improves the safety and reliability of trajectory planning while enhancing driving comfort and reducing unnecessary acceleration, deceleration, or lane changes.

[0084] In some embodiments, the target trajectory planning model includes a multimodal large language model, a multi-stage reasoning process, and a target vehicle trajectory planning expression form including a text form.

[0085] However, the application of existing large language models in closed-loop end-to-end autonomous driving tasks has the following shortcomings:

[0086] First, the modality of supervision is single, relying mostly on a single text modality as a supervisory signal. It fails to effectively model the collaborative relationship between multimodal data such as images, semantic graphs, and driving states, which limits its ability to understand and perceive complex traffic situations.

[0087] Second, there is a lack of an explicit reasoning chain mechanism. Although some studies have tried interactive fine-tuning strategies such as multi-round question-answering, they have not fully mobilized the chain thinking reasoning mechanism, and have performed poorly in complex task decomposition and causal reasoning.

[0088] Third, there is a lack of high-quality planning-layer reasoning datasets. Currently, there is a lack of high-quality datasets that cover real traffic scenarios and have clear reasoning objectives, which makes it difficult to effectively support training and evaluation at the path planning and decision-making control levels, thereby limiting its performance beyond that of traditional imitation learning methods.

[0089] Through the pre-training knowledge of the large language model in this application, the model can integrate visual features and instruction data in text form to generate an explainable multi-stage reasoning process and output executable trajectory planning in text form, effectively solving the problem of missing reasoning chain. It can significantly enhance the hierarchical and explainable reasoning ability of the model in the decision-making process, and improve its behavioral stability and generalization ability in complex environments.

[0090] In addition, the large language model is used to perform scene understanding, traffic sign recognition, key target identification and risk assessment, meta-action decision-making and other stages in the multi-stage reasoning process, fully leveraging the common sense reasoning ability of the large language model, making it more suitable for decision-making and planning tasks in autonomous driving, thereby improving the reasoning ability and planning performance of the trajectory planning model in closed-loop end-to-end autonomous driving.

[0091] Finally, by combining the large language model in this application with multimodal data (including image data and text data), the model's multimodal context understanding ability and structured reasoning ability are significantly enhanced, enabling the model to better perceive environmental changes and make inferences and decisions based on contextual features, thereby improving the model's path planning performance and distributed generalization capabilities in closed-loop scenarios.

[0092] In some embodiments, this application also provides a large-scale, high-quality instruction reasoning dataset for the training of large language models. This dataset includes a large number of training samples and a large number of test samples, comprehensively covering the key decision-making processes in typical autonomous driving scenarios. For example, this dataset, based on Bench2Drive simulation data, records chained reasoning information for the following decision-making stages through automated annotation combined with manual review: scene semantic parsing; traffic rules and sign recognition; key goal and potential risk assessment; and high-level decision action generation.

[0093] To ensure the accuracy and availability of the dataset, the datasets used in the reasoning chain of this application have been manually verified and have good structural consistency and semantic rationality, which effectively improves the model's path planning performance and distribution generalization capabilities in closed-loop scenarios.

[0094] Figure 2 FIG. 1 is a flow chart of a trajectory planning method according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps.

[0095] Step S210: Acquire visual data and command data of the current scene collected by the target vehicle.

[0096] In step S220 , the trajectory planning model is used to process the visual data and command data of the current scene to obtain a multi-stage reasoning process and target vehicle trajectory planning.

[0097] Specifically, the trajectory planning model is obtained based on the training method described in the aforementioned embodiment.

[0098] Optionally, the system first collects visual data and command data from the target vehicle regarding the current scene. The visual data includes camera and lidar information, while the command data includes navigation and trajectory prediction information. The system then uses a trained trajectory planning model, integrating multi-stage reasoning, to analyze the scene and predict the trajectory. For example, at an intersection, the model combines visual data and navigation commands to generate a left-turn trajectory plan, ensuring safe and efficient driving.

[0099] The trajectory planning model in this application can adapt to complex and changeable traffic conditions more accurately, significantly improving the efficiency and accuracy of end-to-end trajectory planning, and enhancing the overall performance and reliability of the autonomous driving system.

[0100] Combined with the above Figure 1 and Figure 2 , describes the method embodiment of the present application in detail, and the following is combined with Figure 3 and Figure 4 , the device embodiment of the present application is described in detail. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment, so for parts not described in detail, reference can be made to the previous method embodiment.

[0101] Figure 3 The figure shows a schematic diagram of the structure of a training device provided in one embodiment of the present application. Figure 3 As shown, the training device 30 provided in the embodiment of the present application includes:

[0102] A first determination module 310 is configured to determine predicted visual data for a next scene based on actual visual data of a current scene;

[0103] A second determination module 320 is configured to determine a first adjustment target parameter based on the predicted visual data of the next scene and the actual visual data of the next scene;

[0104] a third determination module 330 for determining a multi-stage reasoning process based on the actual visual data of the current scene, the instruction data, and the predicted visual data of the next scene, and generating a target vehicle trajectory plan based on the multi-stage reasoning process and the predicted visual data of the next scene;

[0105] A fourth determination module 340 is configured to determine a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer for the next scenario;

[0106] The adjustment module 350 is configured to adjust the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter.

[0107] In one embodiment of the present application, the second determination module 320 is further used to determine feature difference data between the predicted visual data of the next scene and the actual visual data of the next scene using a mean square error loss function; and determine the feature difference data as the first adjustment target parameter.

[0108] In one embodiment of the present application, the fourth determination module 340 is also used to determine the cross entropy loss based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer of the next scene using the cross entropy loss function; and determine the cross entropy loss as the second adjustment target parameter.

[0109] In one embodiment of the present application, the trajectory planning model includes a scene prediction module and a trajectory planning module, and the adjustment module 350 is further used to adjust the scene prediction module and the trajectory planning module to minimize the sum of the first adjustment target parameter and the second adjustment target parameter.

[0110] In one embodiment of the present application, the actual visual data of the current scene includes an actual front view image, and the predicted visual data of the next scene includes a predicted front view image. The first determination module 310 is also used to divide the actual front view image of the current scene into multiple grid images; based on the multiple grid images and the actual front view image of the current scene, determine the predicted front view image of the next scene.

[0111] In one embodiment of the present application, the first determination module 310 is also used to extract features from the actual visual data of the current scene to obtain visual features; determine the driving information and navigation instructions of the target vehicle, and determine the context features corresponding to the driving information and navigation instructions; fuse the visual features and the context features to obtain a first fused feature; and determine the predicted visual data of the next scene based on the first fused feature.

[0112] In one embodiment of the present application, the third determination module 330 is also used to fuse the actual visual data and instruction data of the current scene to generate a second fusion feature, where the instruction data includes navigation instruction data and trajectory prediction instruction data; based on the second fusion feature and the predicted visual data of the next scene, determine a multi-stage reasoning process, where the multi-stage reasoning process includes at least one of scene understanding, traffic sign recognition, target obstacle recognition, and driving risk assessment.

[0113] In one embodiment of the present application, the instruction data includes instruction data in text form, and the third determination module 330 is further used to determine multiple text units corresponding to the instruction data, and visual placeholder symbols contained in the multiple text units; replace the visual placeholder symbols with actual visual data of the current scene associated with the semantic features of the target text unit to generate a second fusion feature, and the target text unit represents the text unit associated with the position of the visual placeholder symbol.

[0114] In one embodiment of the present application, the multi-stage reasoning process includes a meta-action decision, and the third determination module 330 is further used to generate a target vehicle trajectory plan based on the meta-action decision and the predicted visual data of the next scene.

[0115] Figure 4 The figure shows a schematic diagram of the structure of a trajectory planning device provided by an embodiment of the present application. Figure 4 As shown, the trajectory planning device 40 provided in the embodiment of the present application includes:

[0116] An acquisition module 410 is used to acquire visual data and instruction data of the current scene collected by the target vehicle;

[0117] The processing module 420 is used to use the trajectory planning model to process the visual data and command data of the current scene to obtain a multi-stage reasoning process and a target vehicle trajectory plan.

[0118] Below, reference Figure 5 To describe the electronic device according to the embodiment of the present application. Figure 5 Shown is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.

[0119] like Figure 5 As shown, the electronic device 50 includes one or more processors 501 and a memory 502 .

[0120] The processor 501 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 50 to perform desired functions.

[0121] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the methods of the various embodiments of the present application described above and / or other desired functions. Various contents such as actual visual data of the current scene, predicted visual data of the next scene, first adjustment target parameters, instruction data, and second adjustment target parameters may also be stored in the computer-readable storage medium.

[0122] In one example, the electronic device 50 may further include an input device 503 and an output device 504 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0123] The input device 503 may include, for example, a keyboard, a mouse, and the like.

[0124] The output device 504 can output various information to the outside, including actual visual data of the current scene, predicted visual data of the next scene, first adjustment target parameters, instruction data, second adjustment target parameters, etc. The output device 504 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0125] Of course, to simplify, Figure 5 Only some of the components related to the present application in the electronic device 50 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 50 may further include any other appropriate components according to specific application scenarios.

[0126] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present application described above in this specification.

[0127] The computer program product may be written in any combination of one or more programming languages ​​to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0128] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the method according to various embodiments of the present application described above in this specification.

[0129] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0130] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.

[0131] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0132] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0133] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0134] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A training method, characterized in that: For training a trajectory planning model, the method includes: Determining predicted visual data for the next scene based on actual visual data for the current scene; determining a first adjustment target parameter according to the predicted visual data of the next scene and the actual visual data of the next scene; Determining a multi-stage reasoning process based on the actual visual data of the current scene, the instruction data, and the predicted visual data of the next scene, and generating a target vehicle trajectory plan based on the multi-stage reasoning process and the predicted visual data of the next scene; Determining a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer for the next scenario; The trajectory planning model is adjusted based on the first adjustment target parameter and the second adjustment target parameter.

2. The training method according to claim 1, characterized in that The determining the first adjustment target parameter according to the predicted visual data of the next scene and the actual visual data of the next scene includes: Determining feature difference data between the predicted visual data of the next scene and the actual visual data of the next scene using a mean square error loss function; The characteristic difference data is determined as the first adjustment target parameter.

3. The training method according to claim 1, characterized in that The determining of a second adjustment target parameter based on the multi-stage reasoning process, the target vehicle trajectory planning, the multi-stage reasoning process answer for the next scenario, and the target vehicle trajectory planning answer includes: Determining a cross entropy loss using a cross entropy loss function based on the multi-stage reasoning process, the target vehicle trajectory planning, and the multi-stage reasoning process answer and the target vehicle trajectory planning answer for the next scenario; The cross entropy loss is determined as the second adjustment target parameter.

4. The training method according to claim 1, characterized in that The trajectory planning model includes a scene prediction module and a trajectory planning module, and adjusting the trajectory planning model based on the first adjustment target parameter and the second adjustment target parameter includes: The scene prediction module and the trajectory planning module are adjusted to minimize the sum of the first adjustment target parameter and the second adjustment target parameter.

5. The training method according to any one of claims 1 to 4, characterized in that: The actual visual data of the current scene includes an actual front view image, the predicted visual data of the next scene includes a predicted front view image, and determining the predicted visual data of the next scene based on the actual visual data of the current scene includes: dividing the actual front view image of the current scene into a plurality of grid images; A predicted front view image of the next scene is determined based on the plurality of grid images and the actual front view image of the current scene.

6. The training method according to any one of claims 1 to 4, characterized in that: The determining of predicted visual data of the next scene based on actual visual data of the current scene includes: Extracting features from the actual visual data of the current scene to obtain visual features; Determining driving information and navigation instructions of the target vehicle, and determining context features corresponding to the driving information and the navigation instructions; Fusing the visual feature and the context feature to obtain a first fused feature; Based on the first fusion feature, predicted visual data of the next scene is determined.

7. The training method according to any one of claims 1 to 4, characterized in that: Determining a multi-stage reasoning process based on the actual visual data of the current scene, the instruction data, and the predicted visual data of the next scene includes: fusing the actual visual data of the current scene with the instruction data to generate a second fused feature, wherein the instruction data includes navigation instruction data and trajectory prediction instruction data; Based on the second fusion feature and the predicted visual data of the next scene, the multi-stage reasoning process is determined, and the multi-stage reasoning process includes at least one of scene understanding, traffic sign recognition, target obstacle recognition, and driving risk assessment.

8. The training method according to claim 7, characterized in that: The instruction data includes instruction data in text form, and fusing the actual visual data of the current scene with the instruction data to generate a second fusion feature includes: Determining a plurality of text units corresponding to the instruction data, and visual placeholder symbols contained in the plurality of text units; The visual placeholder is replaced with actual visual data of the current scene associated with a semantic feature of a target text unit to generate the second fused feature, where the target text unit represents a text unit associated with the position of the visual placeholder.

9. The training method according to any one of claims 1 to 4, characterized in that: The multi-stage reasoning process includes meta-action decision making, and generating target vehicle trajectory planning based on the multi-stage reasoning process and the predicted visual data of the next scene includes: The target vehicle trajectory plan is generated based on the meta-action decision and the predicted visual data of the next scene.

10. A trajectory planning method, characterized in that: include: Obtain visual data and command data of the current scene collected by the target vehicle; The trajectory planning model is used to process the visual data of the current scene and the instruction data to obtain a multi-stage reasoning process and target vehicle trajectory planning, wherein the trajectory planning model is trained based on the method described in any one of claims 1 to 9.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is used to execute the training method described in any one of claims 1 to 9 and / or the trajectory planning method described in claim 10.

Citation Information

Patent Citations

  • Training method of trajectory planning model, vehicle control mode and device

    CN116859931A

  • Visual language navigation method combining image description and text generation image

    CN117571014A

  • Automatic driving vehicle and track planning and model training method, device and equipment

    CN117746360A

  • Automatic driving vehicle trajectory planning method

    CN119283896A

  • Automatic driving model based on multi-modal large model, training method and automatic driving method

    CN119514635A

Cited By

  • Trajectory prediction system and method based on multi-modal thinking chain

    CN120873507A