Multi-modal path planning and prediction model training method, multi-modal path planning and prediction application method, electronic equipment, computer readable storage medium and computer program product

Through multimodal path planning and prediction models, independent encoders and decoders are trained using multi-dimensional data to generate planning trajectories of bicycles and other vehicles, solving the problem of insufficient stability and flexibility of multimodal path planning in the existing technology, and achieving better autonomous driving effects.

CN120257086APending Publication Date: 2025-07-04MUSHROOM CHELIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312724.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing autonomous driving path planning methods are difficult to generate clear multimodal path planning results. Especially in complex and dynamic scenarios, deep learning-based methods have ambiguity in data training, resulting in insufficient stability and flexibility of path planning.

Method used

The multimodal path planning and prediction model is adopted. By obtaining the original training data of multi-dimensionality, inputting the independent encoder separately to generate encoding features, integrating features with interactive encoder, and generating the bicycle planning trajectory and other vehicle prediction trajectory through independent decoder. The loss is calculated by combining the real trajectory and modal labels, and the model parameters are updated, and a lightweight and generalized multimodal path planning model is trained.

Benefits of technology

It realizes the generation of multimodal path planning results without adjusting the data distribution, improves the flexibility and safety of autonomous driving, provides rich planning information, and improves the autonomous driving capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257086A_ABST
    Figure CN120257086A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a multi-modal path planning and prediction model, a multi-modal path planning and prediction application method, electronic equipment, a computer readable storage medium and a computer program product, and the training method comprises the steps: obtaining multi-dimensional original training data, and inputting the data into each independent encoder of the model, coding features of multiple dimensions are obtained; inputting the coding features of the plurality of dimensions into an interactive encoder to obtain interactive coding features; inputting the interactive coding features into a plurality of independent decoders to obtain own vehicle planning trajectories and other vehicle prediction trajectories of a plurality of driving modes; and calculating loss according to the multi-dimensional decoding result, the real driving track and the driving mode label, and updating model parameters according to the loss. The multi-modal path planning and prediction model trained by the method is simple in structure, light in weight and high in generalization; the model not only can output a multi-mode self-vehicle planning track, but also can predict a multi-mode other-vehicle track, and provides more abundant reference information for subsequent modules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of path planning, and in particular, to a training method for a multi-modal path planning and prediction model, a multi-modal path planning and prediction application method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Currently, there are mainly two implementation schemes for the autonomous driving path planning module: rule-based path planning and deep learning-based path planning. The rule-based method uses predefined rules (such as speed limits, driving rules, and obstacle avoidance requirements) to find the optimal path, which is suitable for fixed environments, simple and reliable but lacking in flexibility, and difficult to handle dynamic and complex scenarios. The deep learning-based method trains a model with human driving data, enabling the vehicle to simulate human decisions in similar scenarios and having better generalization ability.

[0003] The deep learning path planning method requires a large amount of data for training, but different drivers may have different strategies in the same scenario, which may lead to ambiguity in the data, thus affecting the stability of path planning. Algorithm personnel hope that the model can generate multiple path schemes (i.e., multi-modal path planning) to adapt to different strategies.

[0004] However, it is difficult to obtain clear multi-modal path planning results directly relying on data training. Currently, the data distribution is mainly adjusted to increase the amount of data for specific strategies. However, with the increase in data and scenario complexity, simply relying on data adjustment can no longer meet the requirements, making multi-modal path planning a major problem in deep learning path planning. Summary of the Invention

[0005] Embodiments of the present application provide a training method for a multi-modal path planning and prediction model, a multi-modal path planning and prediction application method, an electronic device, a computer-readable storage medium, and a computer program product, so as to simplify the structure of the multi-modal path planning and prediction model and improve the accuracy of multi-modal path planning and prediction.

[0006] Embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, embodiments of the present application provide a training method for a multi-modal path planning and prediction model, wherein the training method for the multi-modal path planning and prediction model includes:

[0008] Obtain multi-dimensional original training data, where the multi-dimensional original training data includes self-vehicle driving data, other-vehicle driving data corresponding to the self-vehicle, and map data;

[0009] Input the original training data in each dimension into the respective independent encoders of the multi-modal path planning and prediction model to obtain encoded features in multiple dimensions. The encoded features in multiple dimensions include the ego-vehicle encoded feature, the other-vehicle encoded feature, and the map encoded feature;

[0010] Input the encoded features in multiple dimensions into the interaction encoder of the multi-modal path planning and prediction model to obtain interaction encoded features;

[0011] Input the interaction encoded features into the respective independent decoders of multiple driving modes of the multi-modal path planning and prediction model to obtain the decoding results of the independent decoders of each driving mode. The decoding results of the independent decoders of each driving mode include the ego-vehicle planned trajectory and the other-vehicle predicted trajectory;

[0012] Calculate the loss of the multi-modal path planning and prediction model based on the decoding results of the independent decoders of each driving mode, the true driving trajectory, and the driving mode label, and update the parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model.

[0013] Optionally, calculating the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders of each driving mode, the true driving trajectory, and the driving mode label includes:

[0014] Calculate the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode by respectively comparing the ego-vehicle planned trajectory and the other-vehicle predicted trajectory output by the independent decoder of each driving mode with the ego-vehicle true trajectory and the other-vehicle true trajectory;

[0015] Generate multiple driving mode masks according to the driving mode label corresponding to the ego-vehicle planned trajectory and the driving mode label corresponding to the other-vehicle predicted trajectory;

[0016] Calculate the loss of the multi-modal path planning and prediction model according to the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the multiple driving mode masks.

[0017] Optionally, calculating the loss of the multi-modal path planning and prediction model according to the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the multiple driving mode masks includes:

[0018] Calculate the loss of each driving mode according to the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the corresponding driving mode mask;

[0019] Calculate the total loss of the multi-modal path planning and prediction model according to the loss of each driving mode.

[0020] Optionally, after obtaining the multi-dimensional original training data, the training method of the multi-modal path planning and prediction model further includes:

[0021] Determine the driving strategy of the host vehicle according to the true driving trajectory of the host vehicle, and mark the driving modality label of the host vehicle according to the driving strategy of the host vehicle;

[0022] Determine the driving strategy of the other vehicle according to the true driving trajectory of the other vehicle, and mark the driving modality label of the other vehicle according to the driving strategy of the other vehicle.

[0023] In a second aspect, an embodiment of the present application further provides a multi-modal path planning and prediction application method, where the multi-modal path planning and prediction application method includes:

[0024] Obtain the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and the map data;

[0025] According to the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and the map data, use the multi-modal path planning and prediction model to output the planned trajectory of the host vehicle in multiple driving modalities and the predicted trajectories of the other vehicles in multiple driving modalities;

[0026] Wherein, the multi-modal path planning and prediction model is trained according to the training method of the multi-modal path planning and prediction model described in any one of the foregoing.

[0027] Optionally, after using the multi-modal path planning and prediction model to output the planned trajectory of the host vehicle in multiple driving modalities and the predicted trajectories of the other vehicles in multiple driving modalities according to the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and the map data, the multi-modal path planning and prediction application method further includes:

[0028] For the planned trajectory of the host vehicle in each driving modality, calculate the collision index between the planned trajectory of the host vehicle in each driving modality and the predicted trajectories of the other vehicles in all driving modalities;

[0029] Determine the target planned trajectory of the host vehicle according to the collision index between the planned trajectory of the host vehicle in each driving modality and the predicted trajectories of the other vehicles in all driving modalities, and the target planned trajectory of the host vehicle is one of the planned trajectories of the host vehicle in multiple driving modalities.

[0030] In a third aspect, an embodiment of the present application further provides a training device for a multi-modal path planning and prediction model, where the training device for the multi-modal path planning and prediction model includes:

[0031] The first acquisition unit is used to acquire multi-dimensional original training data, and the multi-dimensional original training data includes host vehicle driving data, driving data of the other vehicle corresponding to the host vehicle, and map data;

[0032] An encoding unit, configured to input the original training data in each dimension into respective independent encoders of a multi-modal path planning and prediction model, to obtain encoded features in multiple dimensions, where the encoded features in multiple dimensions include ego vehicle encoded features, other vehicle encoded features, and map encoded features;

[0033] An interaction unit, configured to input the encoded features in multiple dimensions into an interaction encoder of the multi-modal path planning and prediction model, to obtain interaction encoded features;

[0034] A decoding unit, configured to input the interaction encoded features into respective independent decoders of multiple driving modes of the multi-modal path planning and prediction model, to obtain decoding results of the respective independent decoders of each driving mode, where the decoding results of the respective independent decoders of each driving mode include an ego vehicle planned trajectory and an other vehicle predicted trajectory;

[0035] An updating unit, configured to calculate a loss of the multi-modal path planning and prediction model according to the decoding results of the respective independent decoders of each driving mode, the true driving trajectory, and the driving mode label, and update parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model, to obtain a trained multi-modal path planning and prediction model.

[0036] In a fourth aspect, an embodiment of the present application further provides a multi-modal path planning and prediction application device, where the multi-modal path planning and prediction application device includes:

[0037] A second obtaining unit, configured to obtain current driving data of an ego vehicle, other vehicle driving data corresponding to the ego vehicle, and map data;

[0038] A planning unit, configured to, according to the current driving data of the ego vehicle, the other vehicle driving data corresponding to the ego vehicle, and the map data, use the multi-modal path planning and prediction model to output an ego vehicle planned trajectory in multiple driving modes and an other vehicle predicted trajectory in multiple driving modes;

[0039] Wherein, the multi-modal path planning and prediction model is trained based on the training device of the multi-modal path planning and prediction model described above.

[0040] In a fifth aspect, an embodiment of the present application further provides an electronic device, including:

[0041] A processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute any one of the foregoing multi-modal path planning and prediction model training methods, and execute any one of the foregoing multi-modal path planning and prediction application methods.

[0042] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium storing one or more programs, which when executed by an electronic device including a plurality of application programs, cause the electronic device to execute any one of the foregoing training methods for the multi-modal path planning and prediction model, and execute any one of the foregoing multi-modal path planning and prediction application methods.

[0043] In a seventh aspect, an embodiment of the present application further provides a computer program product including a computer program or instruction, which when executed by a processor, implements any one of the foregoing training methods for the multi-modal path planning and prediction model, and any one of the foregoing multi-modal path planning and prediction application methods.

[0044] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects: In the training method of the multi-modal path planning and prediction model in the embodiment of the present application, first, multi-dimensional original training data is obtained. The multi-dimensional original training data includes the own vehicle driving data, the other vehicle driving data corresponding to the own vehicle, and the map data. Then, the original training data of each dimension is respectively input into the respective independent encoders of the multi-modal path planning and prediction model to obtain encoded features of multiple dimensions, and the encoded features of multiple dimensions include the own vehicle encoded feature, the other vehicle encoded feature, and the map encoded feature. After that, the encoded features of multiple dimensions are input into the interactive encoder of the multi-modal path planning and prediction model to obtain an interactive encoded feature. Then, the interactive encoded feature is input into the respective independent decoders of multiple driving modes of the multi-modal path planning and prediction model to obtain the decoding results of the respective independent decoders of each driving mode. The decoding results of the respective independent decoders of each driving mode include the own vehicle planned trajectory and the other vehicle predicted trajectory. Finally, the loss of the multi-modal path planning and prediction model is calculated according to the decoding results of the respective independent decoders of each driving mode, the real driving trajectory, and the driving mode label, and the parameters of the multi-modal path planning and prediction model are updated according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model. The training method of the multi-modal path planning and prediction model in the embodiment of the present application trains a lightweight model for vehicle multi-modal path planning and prediction. The model has a simple structure and does not require a large amount of computing resources for model training, and a model with better generalization performance than traditional path planning methods can be obtained. Through the customized multi-modal label, the independent multi-modal trajectory decoder can be trained targeted, and the true multi-modal planning effect of autonomous driving can be achieved without adjusting the data distribution, thereby providing more sufficient planning information for the downstream and improving the autonomous driving ability. In addition, the multi-modal path planning and prediction model trained in the present application can not only output the multi-modal own vehicle planned trajectory, but also predict the multi-modal surrounding vehicle trajectories, providing richer reference information for the subsequent modules, thereby ensuring the driving safety of the own vehicle. Description of the Drawings

[0045] The drawings described herein are provided to further understand the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0046] Figure 1 is a schematic flowchart of a method for training a multi-modal path planning and prediction model in an embodiment of the present application;

[0047] Figure 2 is a schematic diagram of the network structure of a multi-modal path planning and prediction model in an embodiment of the present application;

[0048] Figure 3 is a schematic diagram of a multi-modal driving scenario in an embodiment of the present application;

[0049] Figure 4 is a schematic flowchart of the overall training process of a multi-modal path planning and prediction model in an embodiment of the present application;

[0050] Figure 5 is a schematic flowchart of a method for applying a multi-modal path planning and prediction in an embodiment of the present application;

[0051] Figure 6 is a schematic diagram of the overall process of applying a multi-modal path planning and prediction in an embodiment of the present application;

[0052] Figure 7 is a schematic diagram of the structure of a device for training a multi-modal path planning and prediction model in an embodiment of the present application;

[0053] Figure 8 is a schematic diagram of the structure of a device for applying a multi-modal path planning and prediction in an embodiment of the present application;

[0054] Figure 9 is a schematic diagram of the structure of an electronic device in an embodiment of the present application. Detailed Description of the Embodiments

[0055] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0056] The following will describe in detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.

[0057] An embodiment of the present application provides a training method for a multi-modal path planning and prediction model, as Figure 1 shown, a flowchart of a training method for a multi-modal path planning and prediction model in an embodiment of the present application is provided. The training method for the multi-modal path planning and prediction model at least includes the following steps S110 to step S150:

[0058] Step S110, obtaining multi-dimensional original training data, where the multi-dimensional original training data includes ego vehicle driving data, other vehicle driving data corresponding to the ego vehicle, and map data.

[0059] Combined Figure 2 , a schematic diagram of the network structure of a multi-modal path planning and prediction model in an embodiment of the present application is provided. When training the multi-modal path planning and prediction model, it is necessary to first obtain multi-dimensional data for training the model, including ego vehicle driving data, other vehicle driving data around the ego vehicle, and map data around the ego vehicle. The ego vehicle driving data and other vehicle driving data around the ego vehicle may include data such as the position, heading angle, size (length and width), speed, and acceleration of the vehicle. The map data mainly may include lane line data in the road, etc.

[0060] Step S120, respectively inputting the original training data of each dimension into each independent encoder of the multi-modal path planning and prediction model to obtain encoded features of multiple dimensions, where the encoded features of multiple dimensions include ego vehicle encoded features, other vehicle encoded features, and map encoded features.

[0061] The multi-modal path planning and prediction model trained in the embodiment of the present application includes multiple independent encoders. The role of the multiple independent encoders is to independently encode the training data of the above different dimensions, so as to obtain encoded features of multiple dimensions, including ego vehicle encoded features, other vehicle encoded features, and map encoded features.

[0062] Step S130, inputting the encoded features of multiple dimensions into the interaction encoder of the multi-modal path planning and prediction model to obtain interaction encoded features.

[0063] The multi-modal path planning and prediction model also includes an interaction encoder, and the encoded features from different encoders are sent into the interaction encoder. The interaction encoder is responsible for integrating these features, considering the interactions and dependencies between them, and generating interaction encoded features containing global context information, which helps the model better understand the complex traffic environment and the mutual influence between vehicles.

[0064] Step S140: Input the interactive coding features into the independent decoders of multiple driving modes of the multi-modal path planning and prediction model to obtain the decoding results of the independent decoders of each driving mode. The decoding results of the independent decoders of each driving mode include the planned trajectory of the host vehicle and the predicted trajectories of other vehicles.

[0065] The multi-modal path planning and prediction model further includes independent decoders for multiple driving modes, that is, each driving mode corresponds to an independent decoder. The independent decoders for each driving mode are respectively used to generate the planned trajectory of the host vehicle and the predicted trajectories of other vehicles in each driving mode according to the interactive coding features.

[0066] The above-mentioned multiple driving modes are pre-defined modes, such as straight driving, turning, lane changing, etc. Specifically, which driving modes are defined can be flexibly set by those skilled in the art according to the actual application scenario, and no specific limitation is made here.

[0067] Step S150: Calculate the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders of each driving mode, the actual driving trajectory, and the driving mode label, and update the parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model.

[0068] Based on the planned trajectories of the host vehicle and the predicted trajectories of other vehicles in multiple driving modes decoded in the foregoing steps, combined with the actually collected vehicle driving trajectories and the marked driving mode labels, calculate the deviation between the vehicle planned trajectory and the actual trajectory, that is, the trajectory loss value. Use this trajectory loss value to iteratively update the parameters of the model until the application requirements are met, and output the finally trained multi-modal path planning and prediction model.

[0069] The training method of the multi-modal path planning and prediction model in the embodiments of the present application trains a lightweight model for vehicle multi-modal path planning and prediction. The model structure is simple, and a model with better generalization performance than traditional path planning methods can be obtained without a large amount of computing resources for model training; through the custom multi-modal labels, the independent multi-modal trajectory decoders can be trained specifically, and the true multi-modal planning effect of autonomous driving can be achieved without adjusting the data distribution, thus providing more sufficient planning information for the downstream and improving the autonomous driving ability. In addition, the multi-modal path planning and prediction model trained in the present application can not only output the multi-modal planned trajectories of the host vehicle, but also predict the multi-modal trajectories of surrounding vehicles, providing richer reference information for subsequent modules, thereby ensuring the driving safety of the host vehicle.

[0070] In some embodiments of the present application, the obtaining of multi-dimensional original training data includes: collecting the driving data of the vehicle; collecting the driving data of other vehicles and map data within a predetermined range from the vehicle according to the driving data of the vehicle itself, where the driving data of other vehicles includes the historical driving data and the current driving data of other vehicles; taking the driving data of the vehicle itself as a reference, normalizing the driving data of other vehicles and the map data to obtain the normalized driving data of other vehicles and the map data.

[0071] When obtaining the original data for training the model, the driving data of the vehicle itself can be collected first, including data such as the current position, heading angle, dimensions (length and width), speed, and acceleration of the vehicle itself. Based on the current position of the vehicle itself, the driving data of M other vehicles closest to the vehicle within a certain range around the vehicle, such as within 100 meters, can be collected at the same time. The driving data of other vehicles can include data such as the position, heading angle, dimensions (length and width), speed, and acceleration of other vehicles. At the same time, the above-mentioned driving data of the same other vehicle a period of time ago, such as 2 seconds ago, can also be obtained, so as to improve the accuracy of the subsequent model's prediction of the trajectory of other vehicles by combining the historical data of other vehicles, and further improve the accuracy of the multi-modal planning of the vehicle itself. The map information is mainly the lane line information within a certain range around the current position of the vehicle itself, such as within 100 meters.

[0072] After the above data is collected, the driving data of other vehicles and the map data can be normalized by using translation and rotation operations with the current position of the vehicle itself as the origin and the heading angle of the vehicle itself as the reference direction, that is, the driving data of other vehicles and the map data are uniformly transformed into the coordinate system with the vehicle itself as the origin, so as to facilitate subsequent model processing.

[0073] In some embodiments of the present application, after obtaining the multi-dimensional original training data, the training method of the multi-modal path planning and prediction model further includes: determining the driving strategy of the vehicle itself according to the true driving trajectory of the vehicle itself, and marking the driving modality label of the vehicle itself according to the driving strategy of the vehicle itself; determining the driving strategy of other vehicles according to the true driving trajectory of other vehicles, and marking the driving modality label of other vehicles according to the driving strategy of other vehicles.

[0074] Based on the foregoing embodiments, in order to realize the prediction of the multi-modal trajectory of other vehicles, the embodiments of the present application not only mark the driving modality label of the vehicle itself, but also mark the driving modality label of other vehicles. Specifically, the driving strategy of the vehicle (including the vehicle itself and other vehicles) in the next N seconds can be deduced according to the trajectory points of the vehicle in the next N seconds. For example, if the vehicle goes straight, it is marked as modality 1; if the vehicle has an operation of changing lanes to the left, it is marked as modality 2; if the vehicle has an operation of changing lanes to the right, it is marked as modality 3. Of course, those skilled in the art can flexibly adjust which driving modalities are specifically defined according to the actual application scenario, and no specific limitation is made here.

[0075] For ease of understanding the above embodiments, as Figure 3 shown, a schematic diagram of a multi-modal driving scenario in an embodiment of the present application is provided. In the Figure 3 scenario shown, there are 3 other vehicle targets around the host vehicle. Based on the above definition of the driving mode, the host vehicle belongs to mode 1, the other vehicle 1 belongs to mode 2, the other vehicle 2 belongs to mode 1, and the other vehicle 3 belongs to mode 3.

[0076] In some embodiments of the present application, the independent encoder includes a host vehicle encoder, an other vehicle encoder, and a map encoder. The obtaining of the encoded features in multiple dimensions by respectively inputting the original training data in each dimension into the respective independent encoders of the multi-modal path planning and prediction model includes: inputting the host vehicle driving data into the host vehicle encoder to obtain host vehicle encoded features; inputting the other vehicle driving data into the other vehicle encoder to obtain other vehicle encoded features; and inputting the map data into the map encoder to obtain map encoded features.

[0077] Continuing to refer to Figure 2 , the independent encoder of the embodiment of the present application is composed of three independent encoders: a host vehicle encoder, an other vehicle encoder, and a map encoder.

[0078] Host vehicle encoder: Specifically designed to independently process and analyze the host vehicle driving data, capable of capturing the dynamic behavior of the host vehicle, such as speed changes, acceleration, steering angle, etc., and encoding this information into a higher-level feature representation, which reflects the current state and possible driving intentions of the host vehicle.

[0079] Other vehicle encoder: Specifically designed to independently process the other vehicle driving data. Since the traffic environment may include multiple other vehicles, the other vehicle encoder can process multiple other vehicle data sources and extract the key driving features of each other vehicle, such as position, speed, acceleration, and relative position and speed relative to the host vehicle, which helps the model understand the behavior and possible driving paths of other vehicles.

[0080] Map encoder: Focuses on processing map data. Map data provides detailed information such as road structure, traffic signs, traffic signals, obstacle positions, etc., which is crucial for path planning. The map encoder can encode this complex map information into a feature representation that is easy for the model to understand and utilize, such as road curvature, lane width, traffic signal status, etc.

[0081] Input the self-vehicle driving data into the self-vehicle encoder, and process it through deep learning algorithms (such as convolutional neural networks, recurrent neural networks, etc.) to extract the self-vehicle encoding features that reflect the driving state and intention of the self-vehicle. Input the other-vehicle driving data into the other-vehicle encoder, and also use deep learning algorithms to extract the key driving features of the other vehicle, which may include the speed, acceleration, position information of the other vehicle, as well as the relative position and speed with respect to the self-vehicle. Input the map data into the map encoder, and extract the key information in the map, such as road structure, traffic signal status, obstacle position, etc., through appropriate algorithms to form map encoding features.

[0082] By designing dedicated independent encoders for different types of original training data and extracting encoding features in multiple dimensions, the efficiency and accuracy of data processing are effectively improved, the generalization ability of the model is enhanced, and the accuracy and safety of path planning are improved.

[0083] In some embodiments of the present application, calculating the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders of each driving modality, the true driving trajectory, and the driving modality label includes: calculating the self-vehicle planning loss and the other-vehicle prediction loss of each driving modality by respectively calculating the deviation between the self-vehicle planned trajectory and the other-vehicle predicted trajectory output by the independent decoder of each driving modality and the self-vehicle true trajectory and the other-vehicle true trajectory; generating multiple driving modality masks according to the driving modality label corresponding to the self-vehicle planned trajectory and the driving modality label corresponding to the other-vehicle predicted trajectory; calculating the loss of the multi-modal path planning and prediction model according to the self-vehicle planning loss and the other-vehicle prediction loss of each driving modality and the multiple driving modality masks.

[0084] The independent decoder of each driving modality in the foregoing embodiments will output the self-vehicle planned trajectory and the other-vehicle predicted trajectory. When calculating the model loss value, the deviation between the self-vehicle planned trajectory in each driving modality and the self-vehicle true trajectory can be calculated respectively to obtain the self-vehicle planning loss of each driving modality, and the deviation between the other-vehicle predicted trajectory in each driving modality and the other-vehicle true trajectory can be calculated to obtain the other-vehicle planning loss of each driving modality.

[0085] Since the decoders corresponding to multiple driving modalities perform decoding processing independently, the loss calculation for the decoding result of each decoder only needs to consider the vehicle trajectory data actually belonging to this driving modality. Therefore, this can be achieved through the method of driving modality masks. The role of the driving modality mask is to ensure that only the trajectory points truly belonging to this modality are used for supervised training. For example, as shown in the foregoing embodiments Figure 3 shown, the true trajectories of the self-vehicle and the other vehicle 2 belong to modality 1. Then, for the supervised training of the modality 1 decoder, only the planned trajectory of the self-vehicle and the predicted trajectory of the other vehicle 2 need to be considered.

[0086] By processing the data of different driving modes through an independent decoder, the model can more accurately capture the trajectory features in each mode, thereby improving the accuracy of the ego-vehicle planned trajectory and the predicted trajectories of other vehicles. By introducing a driving mode mask, different mode decoders can be trained specifically for different modes, improving the training accuracy and efficiency of the model.

[0087] In some embodiments of the present application, calculating the loss of the multi-modal path planning and prediction model based on the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode and multiple driving mode masks includes: calculating the loss of each driving mode according to the ego-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the corresponding driving mode mask; calculating the total loss of the multi-modal path planning and prediction model according to the losses of each driving mode.

[0088] Taking Figure 3 the driving scenario shown as an example, the multiple driving mode masks generated in the embodiments of the present application can be respectively expressed as m1_mask = [1, 0, 1, 0], m2_mask = [0, 1, 0, 0], m3_mask = [0, 0, 0, 1]. m1_mask = [1, 0, 1, 0] means that in driving mode 1, only the trajectory errors of the ego-vehicle and other vehicle 2 are considered. m2_mask = [0, 1, 0, 0] means that in driving mode 2, only the trajectory error of other vehicle 1 is considered. m3_mask = [0, 0, 0, 1] means that in driving mode 3, only the trajectory error of other vehicle 3 is considered.

[0089] Based on the above driving state masks, multiply them respectively with the trajectory losses of the corresponding modes, and finally add up the losses of all modes to obtain the final total loss Loss, which can be expressed as: Loss = m1_loss * m1_mask + m2_loss * m2_mask + m3_loss * m3_mask.

[0090] The embodiments of the present application comprehensively consider the losses of multiple driving modes, making the model more robust in the face of complex traffic environments and improving the accuracy of multi-modal path planning and prediction.

[0091] For ease of understanding of the above embodiments, as Figure 4 shown, a schematic diagram of the overall training process of a multi-modal path planning and prediction model in the embodiments of the present application is provided.

[0092] The embodiments of the present application also provide a multi-modal path planning and prediction application method. As Figure 5 shown, a schematic diagram of the process of a multi-modal path planning and prediction application method in the embodiments of the present application is provided. The multi-modal path planning and prediction application method at least includes the following steps S510 to step S520:

[0093] Step S510, obtain the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data;

[0094] Step S520, according to the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data, use the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modes and the predicted trajectories of other vehicles in multiple driving modes;

[0095] Wherein, the multi-modal path planning and prediction model is trained based on the training method of the multi-modal path planning and prediction model described in any one of the foregoing.

[0096] Combined Figure 6 , a schematic diagram of the overall process of a multi-modal path planning and prediction application in an embodiment of the present application is provided. When performing multi-modal path planning and prediction, the trained multi-modal path planning and prediction model described in the foregoing embodiment can be applied. After preprocessing the currently collected driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data, input them into the trained multi-modal path planning and prediction model, and the model can directly output the planned trajectories of the host vehicle in multiple driving modes and the predicted trajectories of other vehicles, meeting the multi-modal trajectory planning requirements in the actual application scenario, and providing richer reference information for subsequent modules, thereby ensuring the driving safety of the host vehicle.

[0097] In some embodiments of the present application, after using the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modes and the predicted trajectories of other vehicles in multiple driving modes according to the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data, the multi-modal path planning and prediction application method further includes: for the planned trajectory of the host vehicle in each driving mode, calculate the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes; according to the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes, determine the target planned trajectory of the host vehicle, and the target planned trajectory of the host vehicle is one of the planned trajectories of the host vehicle in multiple driving modes.

[0098] Continue to refer to Figure 6 , in autonomous driving, the control module will convert the trajectory points of the host vehicle at future moments output by the path planning module into control parameters to control the host vehicle to drive forward. And the multi-modal planning scheme designed in the foregoing embodiment will output multiple host vehicle trajectories at N future moments at each moment. Therefore, a trajectory screening scheme can be designed, and based on this trajectory screening scheme, a clear trajectory can be specified before transmitting the planned trajectory to the control module.

[0099] One implementation of the above trajectory screening scheme can be to calculate the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes respectively, and perform trajectory screening based on the magnitude of the collision index. For example, it may include:

[0100] 1) For the planned trajectory points of the host vehicle in Mode 1, calculate the collision index with all the predicted trajectory points of other vehicles in the three driving modes respectively;

[0101] 2) For the planned trajectory points of the host vehicle in Mode 2, calculate the collision index with all the predicted trajectory points of other vehicles in the three driving modes respectively;

[0102] 3) For the planned trajectory points of the host vehicle in Mode 3, calculate the collision index with all the predicted trajectory points of other vehicles in the three driving modes respectively.

[0103] For the specific calculation of the collision index, first calculate the four corner points of the host vehicle and other vehicles according to the positions, shapes (length and width), and heading angles of the host vehicle and other vehicles at each moment, form two quadrilaterals, and then calculate the Intersection over Union (IoU) of these two quadrilaterals. The sum of the IoUs at all moments is the collision index of the planned trajectory of the host vehicle and the predicted trajectory of other vehicles.

[0104] According to the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes, select the one with the smallest collision index as the target planned trajectory of the host vehicle, which means that the possibility of this trajectory colliding with other vehicles in the planned future time period is the lowest.

[0105] The multi-modal path planning and prediction application method of the embodiments of the present application can consider different driving strategies and scenarios, enabling the autonomous driving vehicle to more flexibly adapt to various road and traffic conditions. By calculating the collision index and selecting the trajectory with the lowest collision possibility as the target trajectory, the safety of the autonomous driving vehicle in a complex traffic environment can be significantly improved. By comprehensively considering the driving data of the host vehicle and other vehicles and map information, more accurate and reasonable driving decisions can be provided for the autonomous driving vehicle, thereby improving the driving efficiency and riding comfort.

[0106] The embodiments of the present application also provide a training device 700 for a multi-modal path planning and prediction model, as Figure 7 shown, which provides a structural schematic diagram of a training device for a multi-modal path planning and prediction model in the embodiments of the present application. The training device 700 for the multi-modal path planning and prediction model includes: a first acquisition unit 710, an encoding unit 720, an interaction unit 730, a decoding unit 740, and an updating unit 750; where:

[0107] A first acquisition unit 710, configured to acquire multi-dimensional original training data, where the multi-dimensional original training data includes self-vehicle driving data, other-vehicle driving data corresponding to the self-vehicle, and map data;

[0108] An encoding unit 720, configured to respectively input the original training data of each dimension into respective independent encoders of a multi-modal path planning and prediction model to obtain encoded features of multiple dimensions, where the encoded features of multiple dimensions include self-vehicle encoded features, other-vehicle encoded features, and map encoded features;

[0109] An interaction unit 730, configured to input the encoded features of multiple dimensions into an interaction encoder of the multi-modal path planning and prediction model to obtain interaction encoded features;

[0110] A decoding unit 740, configured to input the interaction encoded features into respective independent decoders of multiple driving modes of the multi-modal path planning and prediction model to obtain decoding results of the respective independent decoders of each driving mode, where the decoding results of the respective independent decoders of each driving mode include a self-vehicle planned trajectory and an other-vehicle predicted trajectory;

[0111] An updating unit 750, configured to calculate a loss of the multi-modal path planning and prediction model according to the decoding results of the respective independent decoders of each driving mode, the true driving trajectory, and the driving mode label, and update parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model.

[0112] In some embodiments of the present application, the updating unit 750 is specifically configured to: calculate a self-vehicle planning loss and an other-vehicle prediction loss of each driving mode by respectively calculating the self-vehicle planned trajectory and the other-vehicle predicted trajectory output by the respective independent decoders of each driving mode with the self-vehicle true trajectory and the other-vehicle true trajectory; generate multiple driving mode masks according to the driving mode label corresponding to the self-vehicle planned trajectory and the driving mode label corresponding to the other-vehicle predicted trajectory; calculate the loss of the multi-modal path planning and prediction model according to the self-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the multiple driving mode masks.

[0113] In some embodiments of the present application, the updating unit 750 is specifically configured to: calculate the loss of each driving mode according to the self-vehicle planning loss and the other-vehicle prediction loss of each driving mode and the corresponding driving mode mask; calculate the total loss of the multi-modal path planning and prediction model according to the loss of each driving mode.

[0114] In some embodiments of the present application, the training device 600 for the multi-modal path planning and prediction model further includes: a first labeling unit, configured to determine the driving strategy of the host vehicle according to the true driving trajectory of the host vehicle after obtaining multi-dimensional original training data, and label the driving modality label of the host vehicle according to the driving strategy of the host vehicle; a second labeling unit, configured to determine the driving strategy of the other vehicle according to the true driving trajectory of the other vehicle, and label the driving modality label of the other vehicle according to the driving strategy of the other vehicle.

[0115] It can be understood that the above-mentioned training device for the multi-modal path planning and prediction model can implement each step of the training method for the multi-modal path planning and prediction model provided in the foregoing embodiments. The relevant explanations regarding the training method for the multi-modal path planning and prediction model are applicable to the training device for the multi-modal path planning and prediction model, and will not be elaborated herein.

[0116] Embodiments of the present application further provide a multi-modal path planning and prediction application device 800, as Figure 8 shown, which provides a structural schematic diagram of a multi-modal path planning and prediction application device in an embodiment of the present application. The multi-modal path planning and prediction application device 800 includes: a second acquisition unit 810 and a planning unit 820; wherein:

[0117] The second acquisition unit 810 is configured to acquire the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and map data;

[0118] The planning unit 820 is configured to use the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modalities and the predicted trajectories of the other vehicle in multiple driving modalities according to the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and the map data;

[0119] Wherein, the multi-modal path planning and prediction model is trained based on the foregoing training device for the multi-modal path planning and prediction model.

[0120] In some embodiments of the present application, the multi-modal path planning and prediction application device 800 further includes: a calculation unit, configured to calculate the collision index between the planned trajectory of the host vehicle in each driving modality and the predicted trajectories of the other vehicle in all driving modalities respectively after using the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modalities and the predicted trajectories of the other vehicle in multiple driving modalities according to the current driving data of the host vehicle, the driving data of the other vehicle corresponding to the host vehicle, and the map data; a screening unit, configured to determine the target planned trajectory of the host vehicle according to the collision index between the planned trajectory of the host vehicle in each driving modality and the predicted trajectories of the other vehicle in all driving modalities, where the target planned trajectory of the host vehicle is one of the planned trajectories of the host vehicle in multiple driving modalities.

[0121] It can be understood that the above multi-modal path planning and prediction application device can implement each step of the multi-modal path planning and prediction application method provided in the foregoing embodiments. The relevant explanations regarding the multi-modal path planning and prediction application method are applicable to the multi-modal path planning and prediction application device, and will not be elaborated here.

[0122] Figure 9 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 9 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0123] The processor, network interface, and memory can be interconnected through an internal bus. The internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only a bidirectional arrow is used in

[0124] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0125] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, and forms a training device for the multi-modal path planning and prediction model and a multi-modal path planning and prediction application device at the logical level. The processor executes the program stored in the memory.

[0126] The above as shown in the present application Figure 1 The method executed by the training device for the multi-modal path planning and prediction model disclosed in the embodiment shown and Figure 5The method executed by the multi-modal path planning and prediction application device disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0127] The embodiments of the present application also propose a computer program product. The computer program product stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 the method executed by the training device of the multi-modal path planning and prediction model in the illustrated embodiment and execute Figure 5 the method executed by the multi-modal path planning and prediction application device in the illustrated embodiment.

[0128] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing device create means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0130] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0131] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0132] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0133] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.

[0134] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0135] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0136] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A training method for a multi-modal path planning and prediction model, wherein, The training method of the multi-modal path planning and prediction model includes: Obtain multi-dimensional original training data, where the multi-dimensional original training data includes the ego vehicle's driving data, the driving data of other vehicles corresponding to the ego vehicle, and map data; Input the original training data of each dimension into the respective independent encoders of the multi-modal path planning and prediction model to obtain encoded features of multiple dimensions, where the encoded features of multiple dimensions include ego vehicle encoded features, other vehicle encoded features, and map encoded features; Input the encoded features of multiple dimensions into the interaction encoder of the multi-modal path planning and prediction model to obtain interaction encoded features; Input the interaction encoded features into the independent decoders of multiple driving modes of the multi-modal path planning and prediction model to obtain the decoding results of the independent decoders of each driving mode, where the decoding results of the independent decoders of each driving mode include the ego vehicle's planned trajectory and the other vehicle's predicted trajectory; Calculate the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders of each driving mode, the true driving trajectory, and the driving mode label, and update the parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model.

2. The training method of the multi-modal path planning and prediction model according to claim 1, wherein, The calculating the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders of each driving mode, the true driving trajectory, and the driving mode label includes: Calculate the ego vehicle's planning loss and the other vehicle's prediction loss of each driving mode by respectively comparing the ego vehicle's planned trajectory and the other vehicle's predicted trajectory output by the independent decoder of each driving mode with the ego vehicle's true trajectory and the other vehicle's true trajectory; Generate multiple driving mode masks according to the driving mode label corresponding to the ego vehicle's planned trajectory and the driving mode label corresponding to the other vehicle's predicted trajectory; Calculate the loss of the multi-modal path planning and prediction model according to the ego vehicle's planning loss and the other vehicle's prediction loss of each driving mode and the multiple driving mode masks.

3. The training method of the multi-modal path planning and prediction model according to claim 2, wherein, The calculating the loss of the multi-modal path planning and prediction model according to the ego vehicle's planning loss and the other vehicle's prediction loss of each driving mode and the multiple driving mode masks includes: Calculate the loss of each driving mode according to the ego vehicle's planning loss and the other vehicle's prediction loss of each driving mode and the corresponding driving mode mask; Calculate the total loss of the multi-modal path planning and prediction model according to the loss of each driving mode.

4. The training method of the multi-modal path planning and prediction model according to any one of claims 1 to 3, wherein, After obtaining the multi-dimensional original training data, the training method of the multi-modal path planning and prediction model further includes: Determine the driving strategy of the ego vehicle according to the true driving trajectory of the ego vehicle, and mark the driving mode label of the ego vehicle according to the driving strategy of the ego vehicle; Determine the driving strategy of the other vehicle according to the true driving trajectory of the other vehicle, and mark the driving mode label of the other vehicle according to the driving strategy of the other vehicle.

5. A multi-modal path planning and prediction application method, wherein, The multi-modal path planning and prediction application method includes: Obtain the ego vehicle's current driving data, the driving data of other vehicles corresponding to the ego vehicle, and map data; According to the ego vehicle's current driving data, the driving data of other vehicles corresponding to the ego vehicle, and the map data, use the multi-modal path planning and prediction model to output the ego vehicle's planned trajectory of multiple driving modes and the other vehicle's predicted trajectory of multiple driving modes; Among them, the multi-modal path planning and prediction model is trained based on the training method of the multi-modal path planning and prediction model according to any one of claims 1 to 4.

6. The multimodal path planning and prediction application method according to claim 5, wherein, After using the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modes and the predicted trajectories of other vehicles in multiple driving modes according to the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data, the multi-modal path planning and prediction application method further includes: For the planned trajectory of the host vehicle in each driving mode, calculate the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes; Determine the target planned trajectory of the host vehicle according to the collision index between the planned trajectory of the host vehicle in each driving mode and the predicted trajectories of other vehicles in all driving modes, where the target planned trajectory of the host vehicle is one of the planned trajectories of the host vehicle in multiple driving modes.

7. A training device for a multi-modal path planning and prediction model, wherein, The training device of the multi-modal path planning and prediction model includes: A first acquisition unit, configured to acquire multi-dimensional original training data, where the multi-dimensional original training data includes the driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data; An encoding unit, configured to respectively input the original training data in each dimension into each independent encoder of the multi-modal path planning and prediction model to obtain encoded features in multiple dimensions, where the encoded features in multiple dimensions include the encoded feature of the host vehicle, the encoded feature of other vehicles, and the encoded feature of the map; An interaction unit, configured to input the encoded features in multiple dimensions into the interaction encoder of the multi-modal path planning and prediction model to obtain interaction encoded features; A decoding unit, configured to input the interaction encoded features into the independent decoders in multiple driving modes of the multi-modal path planning and prediction model to obtain the decoding results of the independent decoders in each driving mode, where the decoding results of the independent decoders in each driving mode include the planned trajectory of the host vehicle and the predicted trajectory of other vehicles; An updating unit, configured to calculate the loss of the multi-modal path planning and prediction model according to the decoding results of the independent decoders in each driving mode, the true driving trajectory, and the driving mode label, and update the parameters of the multi-modal path planning and prediction model according to the loss of the multi-modal path planning and prediction model to obtain a trained multi-modal path planning and prediction model.

8. A multi-modal path planning and prediction application device, wherein, The multi-modal path planning and prediction application device includes: A second acquisition unit, configured to acquire the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data; A planning unit, configured to use the multi-modal path planning and prediction model to output the planned trajectories of the host vehicle in multiple driving modes and the predicted trajectories of other vehicles in multiple driving modes according to the current driving data of the host vehicle, the driving data of other vehicles corresponding to the host vehicle, and the map data; Among them, the multi-modal path planning and prediction model is trained based on the training device of the multi-modal path planning and prediction model according to claim 7.

9. An electronic device, including: A processor; and a memory arranged to store computer-executable instructions that, when executed, cause the processor to execute the training method of any one of claims 1 to 4 of the multimodal path planning and prediction model, and to execute the multimodal path planning and prediction application method of any one of claims 5 to 6.

10. A computer-readable storage medium storing one or more programs that, when executed by an electronic device including a plurality of application programs, cause the electronic device to execute the training method of any one of claims 1 to 4 of the multimodal path planning and prediction model, and to execute the multimodal path planning and prediction application method of any one of claims 5 to 6.

11. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the training method of any one of claims 1 to 4 of the multimodal path planning and prediction model, and the multimodal path planning and prediction application method of any one of claims 5 to 6.