Track determination method and model training method
By using trajectory determination methods in the autonomous driving system, the world model, trajectory selection model and trajectory generation model are used to accurately determine the future trajectory of the autonomous driving equipment, solving the problem of inaccurate future trajectory in the existing technology, and improving the safety and reliability of autonomous driving.
Patent Information
- Application Number
- CN202510214370.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The future trajectory determined by trajectory planning in the prior art may have inaccurate problems, resulting in serious negative impacts during autonomous driving.
It provides a trajectory determination method, by determining the trajectory prediction time range and time step of the target object, obtaining the current state of the target object, and using the world model, trajectory selection model and trajectory generation model for iterative processing, obtaining the action and state of the target object at all time steps, and finally determining the accurate future trajectory.
This method can accurately determine the future trajectory of the autonomous driving equipment and reduce the negative impact of the inaccurate future trajectory in the process of autonomous driving.
Smart Images

Figure CN120063310A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and particularly to a method for determining a trajectory. One or more embodiments of this specification also relate to a model training method, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the continuous development of autonomous driving technology, autonomous driving has gradually been applied in daily life; in autonomous driving technology, trajectory planning is a relatively important ability, and through the trajectory planning ability, the future movement trajectory of the target object can be calculated; however, the future trajectory determined by the current trajectory planning may have inaccurate problems, and the inaccurate future trajectory may cause serious negative impacts during the autonomous driving process. Therefore, how to accurately determine the future trajectory has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] In view of this, the embodiments of this specification provide a method for determining a trajectory. One or more embodiments of this specification also relate to a model training method, a trajectory determination device, a model training device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0004] According to the first aspect of the embodiments of this specification, a method for determining a trajectory is provided, including:
[0005] Determine the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range;
[0006] Obtain the state of the target object at the current time step, where the state includes the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, and the current time step starts from 0;
[0007] According to the state at the current time step and the preset action modality, use the world model, the trajectory selection model, and the trajectory generation model to iteratively obtain the actions and states of the target object at all time steps;
[0008] Determine the future trajectory of the target object according to the actions and states of the target object at all time steps.
[0009] According to the second aspect of the embodiments of this specification, a trajectory determination device is provided, including:
[0010] A time determination module configured to determine the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range;
[0011] A state acquisition module, configured to acquire the state of the target object at the current time step, where the state includes the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, and the current time step starts from 0;
[0012] An action acquisition module, configured to iteratively acquire the actions and states of the target object at all time steps according to the state at the current time step and a preset action modality, using a world model, a trajectory selection model, and a trajectory generation model;
[0013] A future trajectory determination module, configured to determine the future trajectory of the target object according to the actions and states of the target object at all time steps.
[0014] According to a third aspect of the embodiments of the present specification, a model training method is provided, including:
[0015] Determine a sample data set, where the sample data set includes multiple sample data in different scenarios, and each sample data includes the map information corresponding to the sample object, the trajectory information and feasible routes of the sample object, and the trajectory information of other sample objects;
[0016] Segment the trajectory information of the sample object and the other sample objects to obtain the historical trajectories and future trajectories of the sample object and the other sample objects;
[0017] Construct a sample state according to the map information, the historical trajectories of the sample object and the other sample objects;
[0018] Determine a modality label according to the target feasible route and target action modality corresponding to the future trajectory of the sample object, where the target feasible route is any one of the feasible routes, and the target action modality is any one of the preset action modalities;
[0019] Train and obtain a world model, a trajectory selection model, and a trajectory generation model according to the sample state, the future trajectories of the sample object and the other sample objects, the feasible routes, the preset action modality, and the modality label.
[0020] According to a fourth aspect of the embodiments of the present specification, a model training device is provided, including:
[0021] A sample determination module, configured to determine a sample data set, where the sample data set includes multiple sample data in different scenarios, and each sample data includes the map information corresponding to the sample object, the trajectory information and feasible routes of the sample object, and the trajectory information of other sample objects;
[0022] A trajectory segmentation module, configured to segment the trajectory information of the sample object and the other sample objects to obtain the historical trajectories and future trajectories of the sample object and the other sample objects;
[0023] A state construction module, configured to construct a sample state according to the map information and the historical trajectories of the sample object and the other sample objects;
[0024] A modality label determination module, configured to determine a modality label according to the target feasible route and the target action modality corresponding to the future trajectory of the sample object, where the target feasible route is any one of the feasible routes, and the target action modality is any one of the preset action modalities;
[0025] A model training module, configured to train and obtain a world model, a trajectory selection model, and a trajectory generation model according to the sample state, the future trajectories of the sample object and the other sample objects, the feasible routes, the preset action modalities, and the modality labels.
[0026] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, including:
[0027] A memory and a processor;
[0028] Wherein, the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the above-mentioned trajectory determination method or model training method are implemented.
[0029] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned trajectory determination method or model training method are implemented.
[0030] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned trajectory determination method or model training method are implemented.
[0031] One or more embodiments of this specification provide a trajectory determination method. In the process of determining the future trajectory, first, the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range can be determined, and the state of the target object at the current time step can be obtained; second, models with three different capabilities, namely the world model, the trajectory selection model, and the trajectory generation model, can be used to perform targeted and accurate iterative processing on the state at the current time step and the preset action modality, so as to iteratively obtain the actions and states of the target object at all time steps. Finally, based on the actions and states of the target object at all time steps, an accurate future trajectory can be determined for the target object, thereby avoiding the problem of inaccurate future trajectories, and further reducing the serious negative impacts caused by inaccurate future trajectories during the process of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic application diagram of a trajectory determination method provided by an embodiment of this specification;
[0033] Figure 2 is a flowchart of a trajectory determination method provided by an embodiment of this specification;
[0034] Figure 3 is a flowchart of a model training method provided by an embodiment of this specification;
[0035] Figure 4 is a flowchart of the processing process of a model training method provided by an embodiment of this specification;
[0036] Figure 5 is a schematic structural diagram of a trajectory determination device provided by an embodiment of this specification;
[0037] Figure 6 is a schematic structural diagram of a model training device provided by an embodiment of this specification;
[0038] Figure 7 is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0040] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any or all possible combinations of one or more of the associated listed items.
[0041] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0042] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0043] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through large-scale pre-training of the large model with unlabeled corpora, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.
[0044] When large models are applied in practice, only a small number of samples are needed to fine-tune the pre-trained model for application to different tasks. Large models can be widely applied in fields such as natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0045] With the continuous development of autonomous driving technology, autonomous driving has gradually been applied in daily life; in autonomous driving technology, trajectory planning is a relatively important ability, and through the trajectory planning ability, the future movement trajectory of the target object can be calculated; however, the future trajectory determined by the current trajectory planning may have inaccurate problems, and inaccurate future trajectories may cause serious negative impacts during the autonomous driving process.
[0046] Specifically, trajectory planning is crucial in autonomous driving technology. It can use the outputs of the perception and trajectory prediction modules to generate the future pose for the ego vehicle. And the controller in autonomous driving technology is used to track the planned trajectory and generate control commands for closed-loop driving. Currently, learning-based trajectory planning has received attention because it can automate algorithm iteration, eliminate cumbersome rule design, and ensure safety and comfort in various real-world scenarios.
[0047] Regarding the trajectory planning ability, this specification provides three solutions. The first solution is to use imitation learning to train a planner based on human expert demonstrations. This solution takes into account that drivers can handle various real-world scenarios, so the driving expertise of drivers and a large amount of collected driving data can be used to train the planner. However, there are two problems with the above first solution; the first problem is that achieving closed-loop operation is the ultimate challenge in autonomous driving. It uses driving-oriented metrics to evaluate the planned path, such as safety, compliance with traffic rules, comfort, and progress, etc. By adopting this evaluation method, a significant gap between the training and testing phases of the planner is revealed. That is to say, in the training process of the first solution, when the system encounters scenarios that do not exist in the training data distribution, the decision of trajectory planning may be poor. The second problem is that imitation learning is particularly vulnerable to distribution shift and causal confounding problems, which occur when the network inadvertently captures incorrect correlations and develops shortcuts based on input information, mainly due to the imitation loss that depends on expert demonstrations. Based on this, the gap between the training and testing of the first solution is still huge.
[0048] The second solution is the reinforcement learning solution. In the field of autonomous driving, reinforcement learning can handle problems in specific scenarios, such as highway driving, lane changing, and unprotected left turns. For the reinforcement learning solution, a policy can be directly learned in the control space, which includes throttle, brake, and steering commands. However, the problem with this second solution is that due to the high execution frequency of control commands, the simulation process may be time-consuming, and the exploration may be inconsistent.
[0049] The third solution proposes a trajectory planner that learns to define actions as the planned trajectory of the ego vehicle, which can expand the exploration space in time. However, there is a trade-off between the trajectory time range and the training performance in this solution. Increasing the trajectory time range results in weaker reactive behavior and a reduction in the amount of data, while a smaller trajectory time range leads to challenges similar to those encountered in the control space.
[0050] It should be noted that the above solutions usually adopt a model-free environment, making it difficult to apply to complex and diverse real-world scenarios existing in large-scale driving datasets.
[0051] Based on this, to solve the above technical problems, in this specification, a trajectory determination method is provided. One or more embodiments of this specification are also related to a model training method, a trajectory determination device, a model training device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0052] Considering that the number of model parameters of the model is huge and the computing resources of the mobile terminal are limited, the trajectory determination method provided in the embodiments of this application can be applied to, for example, Figure 1 the application scenario shown, but not limited to this. In the application scenario shown in Figure 1 the model is deployed in the server 10, and the server 10 can connect to one or more autonomous driving devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the autonomous driving devices 20 can include, but are not limited to: cars, trains, airplanes, ships, intelligent robots, etc. The autonomous driving devices 20 can interact with the server 10 to implement the invocation of the model, and further implement the method provided in the embodiments of this specification.
[0053] In the embodiments of this specification, the system composed of the autonomous driving device 20 and the server 10 can perform the following steps: The autonomous driving device determines to send the current state of the autonomous driving device to the server 10; the server 10 first performs operations to determine the prediction time range and the corresponding time step, and obtains the state of the autonomous driving device at the current time step; then, using the world model, the trajectory selection model, and the trajectory generation model, based on the state at the current time step and the preset action modality, it performs trajectory prediction to obtain the future trajectory of the autonomous driving device; after obtaining the future trajectory, the server 10 sends the future trajectory to the autonomous driving device 10, so that the autonomous driving device 10 can drive safely and comfortably based on this future trajectory.
[0054] It should be noted that, in the case where the operating resources of the autonomous driving device 10 can meet the deployment and operation conditions of multiple models, the embodiments of this application can be carried out in the autonomous driving device 10.
[0055] See Figure 2 , Figure 2 which shows a flowchart of a trajectory determination method provided according to an embodiment of this specification, specifically including the following steps.
[0056] Step 202: Determine the trajectory prediction time range of the target object and the time step corresponding to the trajectory prediction time range.
[0057] Among them, the target object can be understood as a moving object that can move based on the future trajectory, and this target object can be a type of mobile device; for example, this target object can be a device such as a car, a train, an airplane, a robot, a ship, etc. that can move based on the future trajectory; this target object can be understood as an autonomous driving device; this target object can also be understood as a device that uses autonomous driving technology to move.
[0058] The trajectory prediction time range can be understood as the time range of the future trajectory that needs to be predicted. For example, this predicted trajectory time range can be a time range from the current time to the next 3 seconds, the next 1 minute, or the next 1 hour, etc.; the time range of this future trajectory corresponds to the trajectory prediction time range.
[0059] A time step can be understood as a time segment corresponding to a trajectory prediction time range; or a time step can be understood as a time point or time range corresponding to each action within a trajectory prediction time range. For example, the trajectory prediction time range can be the time range from the current time to 3 seconds in the future. In this method, the time range or time segment corresponding to every 1 second within this trajectory prediction time range can be taken as 1 time step. Therefore, this trajectory prediction time range corresponds to 3 time steps. Or, in this method, the time range or time segment corresponding to every 0.5 seconds within this trajectory prediction time range can be taken as 1 time step. Therefore, this trajectory prediction time range corresponds to 6 time steps.
[0060] Specifically, this method can determine the trajectory prediction time range of the target object and, based on a preset time step size (such as 1 second or 0.5 seconds), determine the time steps corresponding to this prediction time range.
[0061] Taking the application of the trajectory determination method provided in this specification in an autonomous driving scenario as an example, this trajectory determination method will be described. Here, the target object can be an autonomous driving vehicle, and the trajectory prediction time range is 3 seconds in the future. Based on this, the host vehicle (i.e., the autonomous driving vehicle) is approaching an intersection and needs to predict its driving trajectory within the next 3 seconds according to the current traffic conditions. Therefore, the trajectory prediction time range can be determined as T (i.e., 3 seconds); the time step size is 1 second. Therefore, the time steps corresponding to this prediction time range are also T (i.e., these 3 time steps: the 0th time step, the 1st time step, and the 2nd time step); where the time step size can be understood as a preset prediction step length.
[0062] Step 204: Obtain the state of the target object at the current time step, where the state includes the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, and the current time step starts from 0.
[0063] Here, the current time step can be understood as any one of the multiple time steps corresponding to the trajectory prediction time range; and this current time step starts from the 0th time step, and multiple rounds of iterative processing are performed for each time step. Therefore, the time step in each round of iterative processing can be understood as the current time step; and then the state of each time step can be determined.
[0064] The state of the current time step can be understood as the movement state information of the target object at the current time step; for example, when the target object is an autonomous driving vehicle, the state of the current time step can be understood as the vehicle driving state information of the autonomous driving vehicle at the current time step.
[0065] The map information corresponding to the target object can be understood as the information of the map where the target object is located; in the case where the target object is an autonomous vehicle, the map information is the building information, road network, traffic lights, etc. where the autonomous vehicle is located. In the case where the target object is a ship, the map information is the river network where the ship is located.
[0066] The trajectory information can be understood as the movement trajectory information of the target object or other objects at present or in the past. For example, the position and speed of a vehicle.
[0067] The environmental information of the target object can be understood as the trajectory information and environmental information (such as feasible routes) of the target object. Among them, the trajectory information of the target object can be understood as the position and speed of the target object; in the case where the target object is an autonomous vehicle, the trajectory information of the target object is the position and speed of the vehicle itself. The environmental information of the target object can be understood as the current environment of the target object; in the case where the target object is an autonomous vehicle, the environmental information can be the feasible lanes of the autonomous vehicle, such as the lane lines at an intersection where the autonomous vehicle is traveling (such as the left-turn lane and straight-through lane at an intersection). In the case where the target object is a ship, the map information is the river position where the ship is located.
[0068] Other objects can be understood as other objects in the current environment except the target object; in the case where the target object is an autonomous vehicle, other objects can be other vehicles, pedestrians, etc. These other objects can be other environmental agents or other traffic participants.
[0069] The trajectory information of other objects can be understood as the movement trajectory information of other objects (such as other traffic participants) at present or in the past. For example, the position and speed of other objects such as oncoming vehicles and / or pedestrians.
[0070] Continuing with the above example, this method can iteratively generate the future trajectory of the vehicle itself through multiple models. In this process, first, the state at the t-th time step (i.e., the current time step) can be determined as the current state S t , and subsequently, the current state S t can be input into the model for iterative processing.
[0071] Among them, the current time step t starts from 0, and the current state S t includes the map information of the vehicle itself, the environmental information of the vehicle itself (such as lane lines such as the left-turn lane and / or straight-through lane at an intersection), the trajectory information of the vehicle itself (such as the position and speed of the vehicle itself), and the agent information of other agents (i.e., the trajectory information such as the current and past trajectories of other traffic participants, such as the position and speed of oncoming vehicles and pedestrians).
[0072] In one or more embodiments provided in this specification, after obtaining the state of the target object at the current time step, the following steps are further included:
[0073] Using a perspective-invariant module, perform data preprocessing on the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, to obtain the preprocessed state of the target object at the current time step.
[0074] Among them, the perspective-invariant module can be understood as a module for performing data preprocessing on the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects. The perspective-invariant module can be a neural network model or a network layer.
[0075] Specifically, this method uses a perspective-invariant module to perform data filtering on the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, to obtain the filtered state of the target object at the current time step.
[0076] Continuing with the above example, this method can use a perspective-invariant module to preprocess the current state S t Perform operations such as selecting the nearest K neighbors and filtering feasible lanes, and convert the preprocessed current state to the ego-vehicle coordinate system to obtain the model input data (i.e., the preprocessed state of the target object at the current time step).
[0077] Specifically, in the process of trajectory planning, this method introduces a perspective-invariant module to preprocess the pattern and / or state before inputting it into the network, thereby eliminating time information. For the current state S t Regarding the map and agent information in it, a perspective-invariant module can be used to select the K nearest neighbors (i.e., other objects) to the current position of the ego-vehicle, and only input this information into the policy model, where K is set to half of the map and agent elements.
[0078] For the feasible lanes (i.e., modalities) that capture lateral behavior, this method can use a perspective-invariant module to filter out the segments where the point closest to the current position of the ego-vehicle is the starting point, and retain K r points, where K r is set to one-fourth of the N r points in each route. Finally, convert the feasible lanes, agents, and map poses (i.e., actions) to the coordinate system of the ego-vehicle at the current time step t. Then, subtract the historical time steps t - H:t from the current time step t to obtain a time step range of -H:0.
[0079] Based on the above content, the role of the perspective-invariant module is: to eliminate time information and convert the state S tConvert it into a format more suitable for the policy model to process. Specifically, it includes: eliminating time information and perspective invariance.
[0080] Eliminating time information: It means converting the time step range from the current time step \(t\) by subtracting the historical time steps \(t - H:t\) to \(-H:0\), thereby eliminating time information.
[0081] Perspective invariance: It means converting the information of the map, agents, and feasible lanes into the current coordinate system of the ego vehicle, so that the policy model can process information centered on the ego vehicle without being affected by the global coordinate system.
[0082] Through this preprocessing method, the policy model can process state information more efficiently and improve its generalization ability in different scenarios.
[0083] Among them, the specific steps of state preprocessing:
[0084] (1) Select the nearest \(K\) neighbors.
[0085] Its purpose is: to reduce the complexity of state \(S\) t and only retain the information most relevant to the ego vehicle. The implementation method is: for map and agent information, select the \(K\) nearest neighbors to the current position of the ego vehicle. \(K\) is set to half of the map and agent elements. For example, for map information, select the 5 nearest lane lines and 5 traffic lights to the ego vehicle. For agent information, select the 5 nearest other vehicles and 5 pedestrians to the ego vehicle.
[0086] (2) Filter the feasible lanes.
[0087] Its purpose is: to capture lateral behavior (such as lane changing) and reduce the complexity of the feasible lanes. The implementation method is: for the feasible lanes, filter out the segments whose nearest point to the current position of the ego vehicle is the starting point. Retain \(K\) points, and \(K\) is set to one-fourth of the \(N\) points in each route. For example, retain the 10 nearest points to the ego vehicle on each feasible lane.
[0088] (3) Convert to the ego vehicle coordinate system.
[0089] Its purpose is: to convert the information of the map, agents, and feasible lanes into the current coordinate system of the ego vehicle to achieve perspective invariance. The implementation method is: convert the feasible lanes, agents, and map poses into the coordinate system of the ego vehicle at the current time step \(t\). Convert the positions and speeds of oncoming vehicles and pedestrians into the coordinate system of the ego vehicle.
[0090] Step 206: According to the state of the current time step and the preset action modality, use the world model, trajectory selection model, and trajectory generation model to iteratively obtain the actions and states of the target object at all time steps.
[0091] Among them, the preset action modality can be understood as a pre-set action modality; this action modality can be understood as the action type of the target object, and this action modality can be defined as the combination of horizontal multi-modal and vertical multi-modal. The horizontal multi-modal and vertical multi-modal are respectively described by the feasible path (such as the feasible lane of the host vehicle) and the average speed. For example, the horizontal multi-modal is the feasible lane of the host vehicle (such as the straight lane, left-turn lane, right-turn lane), and the vertical multi-modal is described by the average speed (such as slowly decelerating, normally decelerating, constant speed, accelerating, rapidly accelerating, etc.).
[0092] The world model can be understood as a model used to predict the trajectory information of other objects in the next time step (i.e., the next time step of the current time step); the trajectory selection model can be understood as a model used to predict the target modality of the target object in the next time step; the trajectory generation model can be understood as a model that predicts the action of the target object at the current time step.
[0093] Among them, the target modality can be understood as the action modality corresponding to the target object in the next time step from multiple action modalities.
[0094] The action at a time step can be understood as the action corresponding to the target object and the time step. For example, this action can be moving forward 5 meters, turning left 5 meters ahead, etc.
[0095] Specifically, this method can determine the preset action modality corresponding to the target object, and use the world model, trajectory selection model, and trajectory generation model to perform iterative processing based on the state at the current time step and the preset action modality, so as to obtain the actions and states of the target object at all time steps.
[0096] In one or more embodiments provided in this specification, the iterative obtaining of the actions and states of the target object at all time steps by using the world model, trajectory selection model, and trajectory generation model according to the state at the current time step and the preset action modality includes steps one to four:
[0097] Step one: Generate the action of the target object at the current time step according to the state at the current time step and the preset action modality, using the world model, trajectory selection model, and trajectory generation model.
[0098] Specifically, this method can use the world model, trajectory selection model, and trajectory generation model to perform the first round of processing of multiple rounds of iterative processing on the state at the current time step and the preset action modality, and obtain the action of the target object at the current time step.
[0099] Continuing with the above example, during the process of iteratively generating the future trajectory of the ego vehicle, the state at the 0th time step (i.e., the current time step) and all modalities (i.e., preset action modalities) can be predicted and processed using the world model, the trajectory selector (i.e., the trajectory selection model), and the trajectory generation model (i.e., the trajectory generator) to obtain the action of the ego vehicle at the 0th time step.
[0100] In one or more embodiments provided in this specification, generating the action of the target object at the current time step according to the state at the current time step and the preset action modalities, using the world model, the trajectory selection model, and the trajectory generation model, includes:
[0101] Input the state at the current time step into the world model to obtain the trajectory information of other objects at the next time step;
[0102] Input the state at the current time step and the preset action modalities into the trajectory selection model to obtain the target modality of the target object at the next time step;
[0103] Input the current time step, the trajectory information of other objects at the next time step, and the target modality into the trajectory generation model to obtain the action of the target object at the current time step.
[0104] Continuing with the above example, during the process of iteratively generating the future trajectory of the ego vehicle, first, the state at the tth time step can be determined as the current state S t Input it into the world model, and use the world model to output the future trajectories of other traffic participants (oncoming vehicles and pedestrians) at the (t + 1)th time step (i.e., the trajectory information of other objects at the next time step); for example, this trajectory information can be: the oncoming vehicle will continue to go straight through the intersection within the next 3 seconds, and the pedestrian will stay on the sidewalk within the next 3 seconds.
[0105] where t starts from 0, and the current state S t includes map information (lane lines at the intersection, such as the left-turn lane and the straight-through lane at the intersection) and agent information (the current and past trajectories of the ego vehicle and other traffic participants, such as the position and speed of the ego vehicle, the position and speed of the oncoming vehicle and the pedestrian).
[0106] Secondly, input the current state S t and all modalities (left-turn lane and low speed, left-turn lane and high speed, straight-through lane and low speed, straight-through lane and high speed) into the trajectory selector to output the optimal modality (i.e., the target modality).
[0107] Specifically, this method can input the current state S t, all modal input trajectory selectors are used to generate multiple modal probabilities, which are the probabilities that the host vehicle will perform actions corresponding to each mode in the next time step. For example, the probability of turning left lane and low speed is 0.4, the probability of turning left lane and high speed is 0.3, the probability of going straight lane and low speed is 0.2, and the probability of going straight lane and high speed is 0.1;
[0108] Select the mode corresponding to the maximum modal probability from the multiple modal probabilities as the preferred mode. For example, the preferred mode is turning left lane and low speed.
[0109] It should be noted that all modes refer to all modes obtained by combining lateral multimodality and longitudinal multimodality. According to the above example, when there are two lateral multimodalities and five longitudinal multimodalities, all modes can be 10 modes.
[0110] Finally, input the current state S t , the preferred mode, and the future trajectories of other traffic participants at the (t + 1)-th time step into the trajectory generator to output the action a of the host vehicle at the t-th time step 0 (i.e., the action at the current time step).
[0111] For example, at time t, the vehicle is located at the coordinate (0, 0), and the action at time t is to move forward 5 meters (which can be understood as stepping on the accelerator). Correspondingly, in this case, the vehicle position at time (t + 1) will be predicted to have moved forward 5 meters relative to the position at time t, that is, the new position may be (0, 5).
[0112] It should be noted that in the scenario of high-frequency trajectory planning, the world model and the trajectory selector (i.e., the trajectory selection model) will be re-invoked during each trajectory planning to generate predictions and select modes based on the latest environmental information. This method can improve the real-time performance, flexibility, and safety of the system, enabling it to cope with complex traffic environments.
[0113] Step 2: When it is determined that the current time step is not the last time step, determine the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of other objects at the next time step.
[0114] Continuing with the above example, according to the action a of the host vehicle at the t-th time step 0 and the future trajectories of other traffic participants at time step (t + 1) (the trajectory information of other objects at the next time step), calculate the state S at the next moment t+1 (i.e., the state at the next time step).
[0115] Among them, the current and past trajectories (i.e., position and speed) of the ego vehicle are updated according to the action a0 at the t-th time step; the current and past trajectories (i.e., position and speed) of other traffic participants are updated according to the future trajectories at time step t + 1.
[0116] In addition, the action a of the ego vehicle at the t-th time step 0 is saved to the action sequence, and the states at the t-th time step and the (t + 1)-th time step are both saved to the state sequence.
[0117] Step 3: Take the next time step as the current time step, and continue to execute the step of obtaining the state of the target object at the current time step.
[0118] Continuing with the above example, after completing the current time step, take the next time step of the current time step as the current time step, and continue to execute the operations in the above Step 1 to Step 2 to obtain the state of the target object at the current time step.
[0119] In one or more embodiments provided in this specification, the step of continuing to execute the step of obtaining the state of the target object at the current time step includes:
[0120] Input the state of the target object at the current time step and the target modality into the trajectory generation model to obtain the action of the target object at the current time step;
[0121] In the case where it is determined that the current time step is not the last time step, determine the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of the other objects at the next time step;
[0122] Take the next time step as the current time step, and continue to execute the step of obtaining the state of the target object at the current time step.
[0123] Continuing with the above example, in the processing method provided in this embodiment, after using the world model, the trajectory selector (i.e., the trajectory selection model), and the trajectory generator (i.e., the trajectory generator) to complete the first iteration process and obtain the state at the 0-th time step, in subsequent iterations, it is no longer necessary to call the world model and the trajectory selector to predict the future trajectories of other traffic participants and the optimal modality at the next time step. Instead, use the state updated in the previous iteration (i.e., the state at the current time step) and the optimal modality, input them into the trajectory generator, and use the trajectory generator to output the action α t+1 ;
[0124] Then continue to use the trajectory generator for iterative processing until all time steps are processed.
[0125] Based on the above, it can be seen that in the subsequent iteration process of this implementation, only the trajectory generation model is used for execution, thereby reducing the computational complexity in the trajectory planning process and enabling the calculation of future trajectories quickly and efficiently.
[0126] Step Four: Until it is determined that the current time step is the last time step, obtain the actions and states of the target object at all time steps.
[0127] Continuing with the above example, when it is determined that the current time step of the current iteration process is the last time step, it is determined that the iteration ends, and the actions and states of the ego vehicle at all time steps are obtained.
[0128] Step 208: Determine the future trajectory of the target object according to the actions and states of the target object at all time steps.
[0129] Specifically, this method can, after determining the actions and states of the target object at all time steps, fuse the actions and states of the target object at all time steps to obtain the future trajectory of the target object. Among them, the time range of this future trajectory can be consistent with the trajectory prediction time range.
[0130] In one or more embodiments provided in this specification, the determining the future trajectory of the target object according to the actions and states of the target object at all time steps includes:
[0131] Perform a sequence combination on the actions and states of the target object at all time steps to obtain the future trajectory of the target object.
[0132] Continuing with the above example, save the actions a 0 of the ego vehicle at all time steps to the action sequence, and save the states at all time steps to the state sequence; and obtain the future trajectory of the ego vehicle according to the ego vehicle action sequence and state sequence at all time steps. Subsequently, corresponding driving instructions can be generated according to the future trajectory of the ego vehicle, so that the ego vehicle moves based on this driving instruction; thus achieving accurate autonomous driving.
[0133] One or more embodiments of this specification provide a trajectory determination method. In the process of determining the future trajectory, first, the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range can be determined, and the state of the target object at the current time step can be obtained. Secondly, models with three different capabilities, namely the world model, the trajectory selection model, and the trajectory generation model, can be used to perform targeted and accurate iterative processing on the state at the current time step and the preset action modality, so as to iteratively obtain the actions and states of the target object at all time steps. Finally, based on the actions and states of the target object at all time steps, an accurate future trajectory can be determined for the target object, thereby avoiding the problem of inaccurate future trajectories and reducing the serious negative impacts caused by inaccurate future trajectories during the autonomous driving process.
[0134] See Figure 3 , Figure 3 shows a flowchart of a model training method provided according to an embodiment of this specification, which specifically includes the following steps.
[0135] Step 302: Determine a sample data set, where the sample data set includes multiple sample data of different scenarios, and each sample data includes the map information corresponding to the sample object, the trajectory information and feasible routes of the sample object, and the trajectory information of other sample objects.
[0136] Among them, the sample data set can be understood as a set of sample data used for model training; the sample data of different scenarios can be understood as the sample data corresponding to different scenarios, and the different scenarios can be understood as the different scenarios where the sample object is located, such as traffic scenarios (intersections, highways, etc.), river channel scenarios, etc.
[0137] The sample object can be understood as the target object used as a training sample, and other sample objects can be understood as other objects used as training samples; for the explanations of terms such as map information, trajectory information, and feasible routes, reference can be made to the explanations of the above-mentioned trajectory determination method, which will not be elaborated here.
[0138] Taking the application of the model training method provided in this specification in the autonomous driving scenario as an example, the model training method is described as follows: First, this method needs to randomly extract a batch of sample data of different scenarios from the expert driving data to form a sample data set. Among them, the sample data set includes multiple sample data, and each sample data consists of map information (such as road network, traffic lights, etc.), the trajectories of all traffic participants (such as the ego vehicle, other environmental intelligent agents (such as oncoming vehicles)), and the feasible lanes of the ego vehicle (such as straight lanes and left-turn lanes).
[0139] Step 304: Segment the trajectory information of the sample object and the other sample objects to obtain the historical trajectories and future trajectories of the sample object and the other sample objects.
[0140] Specifically, this method segments the trajectory information of the sample object and the trajectory information of the sample object respectively to obtain the historical trajectory and future trajectory of the sample object, and the historical trajectory and future trajectory of the other sample objects; among them, the time range of the historical trajectory of the sample object is the same as the time range of the historical trajectory of the other sample objects, that is to say, the trajectory lengths of the historical trajectories of the sample object and the other sample objects are the same; among them, the time range of the future trajectory of the sample object is the same as the time range of the future trajectory of the other sample objects, that is to say, the trajectory lengths of the future trajectories of the sample object and the other sample objects are the same.
[0141] In one or more embodiments provided in this specification, the segmenting the trajectory information of the sample object and the other sample objects to obtain the historical trajectories and future trajectories of the sample object and the other sample objects includes:
[0142] Segment the trajectory information of the sample object and the other sample objects according to a preset segmentation time point to obtain the historical trajectory and future trajectory of the sample object, and the historical trajectory and future trajectory of the other sample objects.
[0143] Among them, the preset segmentation time point can be understood as a time point set in advance for segmenting the trajectory information. This preset segmentation time point can be the 1st second, the 8th second, the 1st minute, etc. of the trajectory information.
[0144] Continuing with the above example, during the process of model training, this method can divide the trajectories of all traffic participants into historical trajectories and future trajectories according to the specified time point (i.e., the preset segmentation time point). For example, if a certain sample is a driving trajectory that continuously travels for 20 seconds, this driving trajectory can be divided at the 8th second as the specified time point, with 0 - 8 seconds as the historical trajectory and 9 - 20 seconds as the future trajectory.
[0145] By using the preset segmentation time point in the above embodiments to segment the trajectory information of the sample object and the other sample objects, historical trajectories and future trajectories can be obtained. Subsequently, the historical trajectories can be used as training samples, the future trajectories can be used as sample labels, and efficient model training can be achieved using the training samples and sample labels.
[0146] Step 306: Construct a sample state according to the map information, and the historical trajectories of the sample object and the other sample objects.
[0147] Among them, the sample state can be understood as the current trajectory state of the sample object and other sample objects during the process of predicting the future trajectory. For example, the current driving conditions of all traffic participants.
[0148] Specifically, this method constructs an initial state based on the map information, the historical trajectories of the sample object and other sample objects. For example, the initial state is constructed using the map information and the historical trajectories of all traffic participants.
[0149] It should be noted that the state S in this method t can include the map and agent information represented in a vectorized representation. Among them, the map information m includes the road network, traffic lights, etc., and is represented by polylines and polygons. Among them, the agent information includes the current and past position information of the ego vehicle and other environmental agents, and is represented by polylines. The index of the ego vehicle is 0, and the index range of other environmental agents is from 1 to N. For each agent i, its history is represented as where H is the historical time range.
[0150] In one or more embodiments provided in this specification, after constructing the sample state, it further includes:
[0151] Using a view-invariant module, perform data preprocessing on the map information, the historical trajectories of the sample object and the other sample objects in the sample state to obtain the preprocessed sample state.
[0152] For the step of using the view-invariant module to perform data preprocessing on the map information, the historical trajectories of the sample object and the other sample objects in the sample state, reference can be made to the step of "using the view-invariant module to perform data preprocessing on the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects" in the above trajectory determination method, which will not be elaborated here.
[0153] Step 308: Determine a modality label according to the target feasible route and the target action modality corresponding to the future trajectory of the sample object, where the target feasible route is any one of the feasible routes, and the target action modality is any one of the preset action modalities.
[0154] Among them, the target feasible route can be understood as the feasible route of the sample object in the future trajectory; the target action modality can be understood as the preset action modality of the sample object in the future trajectory; it should be noted that this method can define the preset action modality during the model training process, and there can be multiple preset action modalities, including: horizontal multi-modal and vertical multi-modal; the horizontal multi-modal can be the feasible path (i.e., the feasible route), such as the vehicle's own feasible lanes (such as the straight lane, left-turn lane, right-turn lane); the vertical multi-modal can be the average speed description (such as slow down, normal deceleration, uniform speed, speed up, rapid speed up, etc.). For example, a preset action modality including horizontal multi-modal and vertical multi-modal can be: driving in the straight lane at a uniform speed.
[0155] The modality label can be the modality corresponding to the sample object in the future trajectory; for example, the modality label can be: the horizontal multi-modal (such as the straight lane) and the vertical multi-modal (such as uniform speed) corresponding to the future trajectory of the vehicle itself.
[0156] Specifically, this method can determine the target feasible route corresponding to the future trajectory of the sample object and the target action modality corresponding to the future trajectory; use the target feasible route to detect the target action modality, and when the feasible route included in the target action modality is consistent with the target feasible route, determine the target action modality as the modality label; thus avoiding the problems of incorrect training data and non-compliance with traffic rules.
[0157] Continuing with the above example, after constructing the initial state based on the historical trajectory and map information, the future trajectory can be used as a label to train the world model, trajectory selector, and trajectory generator. And, define the modality as the combination of horizontal multi-modal and vertical multi-modal, where the horizontal and vertical multi-modal are respectively described by the vehicle's own feasible lanes and average speed, and define the label modality as the modality where the future label trajectory of the vehicle itself is located.
[0158] Step 310: Train to obtain a world model, a trajectory selection model, and a trajectory generation model according to the sample state, the future trajectories of the sample object and the other sample objects, the feasible route, the preset action modality, and the modality label.
[0159] Specifically, this method uses the sample state, the future trajectories of the sample object and the other sample objects, the feasible route, the preset action modality, and the modality label as training data to perform model training on the world model, the trajectory selection model, and the trajectory generation model, and obtains the trained world model, trajectory selection model, and trajectory generation model.
[0160] It should be noted that in this method, the trajectory planning task can be modeled as a sequential decision-making process, and the autoregressive model can be decoupled into a policy model (i.e., the trajectory generation model) and a state transition model (i.e., the trajectory selection model). The key to connecting trajectory planning and the autoregressive model lies in defining the action as the next pose of the ego vehicle (e.g., the action), that is Therefore, after passing the autoregressive model forward, the decoded poses will be collected as the trajectory planned by the ego vehicle. With this definition and the vectorized representation, the state-action sequence is simplified to the state sequence P:
[0161]
[0162] where S in the above formula is the state, a is the action, m is the map information, H is the historical time range. T is the time range of the future trajectory to be predicted, and N is other environmental agents (i.e., other sample objects).
[0163] Under this definition, the state sequence can be further expressed in an autoregressive manner and decomposed into a policy model and a world model:
[0164]
[0165] Furthermore, a consistent modal information C (i.e., the target action modality) is introduced, which remains unchanged between time steps and is introduced into the autoregressive process:
[0166]
[0167] From this process, the consistent autoregressive model reveals a generation-selection framework, where the trajectory selector scores each mode based on the initial state S 0 while the trajectory generator generates multi-modal trajectories by sampling from the policy. In addition, since the world model in this formula generates the poses of traffic participants at time step t + 1 based on the current state S t it needs to be used at each time step, and this process is very time-consuming. Therefore, a trajectory predictor is used as a non-reactive world model, which generates the future poses of all traffic participants at once based on the initial state S 0
[0168] In one or more embodiments provided in this specification, the world model, the trajectory selection model, and the trajectory generation model are trained according to the sample state, the sample object, the future trajectories of the other sample objects, the feasible routes, and the modal labels, steps one to three:
[0169] Step one: Train and obtain the world model according to the sample state.
[0170] Specifically, this method inputs the sample state into the world model for processing, obtains the predicted future trajectories of other sample objects, and trains the world model based on the predicted future trajectories of other sample objects and the sample labels corresponding to the world model (such as the future trajectories of other sample objects) to obtain the trained world model.
[0171] In one or more embodiments provided in this specification, training the world model according to the sample state includes:
[0172] Input the sample state into the initial world model to obtain the predicted future trajectories of the other sample objects;
[0173] Calculate the first loss value according to the predicted future trajectories of the other sample objects and the future trajectories of the other sample objects;
[0174] According to the first loss value, use the optimizer to update the network parameter weights of the initial world model until the preset iteration end condition is met, and train to obtain the world model.
[0175] The first loss value can be understood as the loss value for adjusting the model parameters of the world model; this first loss value can be set according to the actual scenario, such as the L1 regression loss value, cross-entropy loss value, etc.
[0176] Continuing with the above example, in the process of training the world model, first, randomly initialize the network parameter weights of the world model, and input the initial state S 0 (i.e., the sample state) into the world model to be trained (i.e., the initial world model) to obtain the predicted future trajectories of other environmental agents output by the world model (excluding the ego vehicle);
[0177] Secondly, calculate the L1 regression loss according to the predicted future trajectories of all environmental agents (i.e., the predicted future trajectories of other sample objects) and the corresponding trajectory labels (the future trajectories of other sample objects) to obtain the L1 regression loss value;
[0178] Specifically, this method can obtain the initial state S 0 and the environmental agent label trajectories (i.e., the future trajectories of other sample objects) from the expert driving dataset D, and calculate the L1 loss L of the world model in S2 by the following formula wm :.
[0179]
[0180] Where is the trajectory output by the world model.
[0181] Finally, according to the L1 regression loss value, use the optimizer to update the weights of the world model, and continuously iterate this process until the preset iteration end condition is met, obtaining the trained world model, and fix the weights of this world model.
[0182] In the above embodiments, the world model is trained by using the sample state (training sample) and the future trajectories of other sample objects (sample labels), obtaining a world model that can accurately predict the future trajectories of other sample objects, facilitating subsequent accurate prediction of the future trajectories of the target object.
[0183] Step 2: Train and obtain the trajectory selection model according to the sample state, the feasible route, the preset action modality, and the modality label.
[0184] Specifically, this method uses the sample state, the feasible route, and the preset action modality as training samples, uses the modality label as the sample label, and trains the trajectory selection model with the training samples and the sample label, thereby obtaining the trained trajectory selection model.
[0185] In one or more embodiments provided in this specification, the training and obtaining of the trajectory selection model according to the sample state, the feasible route, the preset action modality, and the modality label includes:
[0186] Determine an action modality set according to the feasible route and the preset action modality;
[0187] Input the sample state and the action modalities in the action modality set into the initial trajectory selection model to obtain the probabilities of the action modalities in the action modality set;
[0188] Calculate a second loss value according to the probabilities of the action modalities in the action modality set and the modality label;
[0189] According to the second loss value, use the optimizer to update the network parameter weights of the initial trajectory selection model until the preset iteration end condition is met, and train and obtain the trajectory selection model.
[0190] Among them, the second loss value can be understood as the loss value for adjusting the model parameters of the trajectory selection model; this second loss value can be set according to the actual scenario, such as the L1 regression loss value, the cross-entropy loss value, etc.
[0191] Continuing with the above example, in the process of training the trajectory selector, first, the network parameter weights of the trajectory selector can be randomly initialized. The initial state and all modalities (in the above example, when there are two lateral multi-modalities and five longitudinal multi-modalities, all modalities are 10 modalities) are input into the trajectory selector (initial trajectory selection model), and the probabilities of each modality output by the trajectory selector (i.e., the probabilities of the action modalities in the action modality set) are obtained.
[0192] Secondly, the cross-entropy loss (i.e., the second loss value) is calculated based on the probabilities of each modality and the label modality (i.e., the modality label). Specifically:
[0193] This method can obtain the initial state s from the expert driving dataset D 0 and each candidate modality c i (i.e., all modalities). The cross-entropy loss L of the trajectory selector is calculated by the following formula selector :
[0194]
[0195] where N mode is the number of candidate patterns, is the indicator function, and σ i is the probability estimate of the trajectory selector for each candidate modality c i .
[0196] Finally, the weights of the trajectory selector are updated using an optimizer, and this process is continuously iterated until the preset iteration end condition is satisfied, obtaining a trained trajectory selector, and fixing the weights of this trajectory selector.
[0197] In the above embodiment, by using the training samples and sample labels to train the trajectory selection model, a trajectory selection model that can accurately predict the probabilities of action modalities is obtained, which is convenient for accurately predicting the future trajectory of the target object subsequently.
[0198] Step three: Train and obtain the trajectory generation model according to the sample state, the modality label, and the world model.
[0199] Among them, the world model is a trained world model.
[0200] In one or more embodiments provided in this specification, the training and obtaining of the trajectory generation model according to the sample state, the modality label, and the world model includes:
[0201] Input the sample state into the world model to obtain the predicted future trajectories of the other sample objects;
[0202] Input the sample status, the modality label, and the predicted future trajectories of the other sample objects into the initial trajectory generation model to obtain the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the sample object;
[0203] Calculate the third loss value, the fourth loss value, and the fifth loss value according to the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the sample object;
[0204] Update the network parameter weights of the initial trajectory generation model using an optimizer according to the third loss value, the fourth loss value, and the fifth loss value until the preset iteration end condition is met, and train to obtain the trajectory generation model.
[0205] Among them, the third loss value, the fourth loss value, and the fifth loss value can be understood as the loss values for adjusting the model parameters of the trajectory generation model; these three loss values can be set according to the actual scenario, such as the MSE loss value, entropy loss value, policy loss value, etc.
[0206] Continuing with the above example, the training method of the trajectory generation model in this method is as follows:
[0207] 1. Input the initial state into the world model (the trained world model) to obtain the predicted future trajectories of other environmental agents output by the world model.
[0208] 2. Input the initial state, label modality, and the predicted future trajectories of other environmental agents into the trajectory generator (i.e., the initial trajectory generation model) to obtain the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the ego vehicle output by the trajectory generator.
[0209] Among them, the state sequence is the state of the ego vehicle within the future time range (such as the above 12 seconds), including information such as position, speed, and direction;
[0210] The action sequence is the actions of the ego vehicle within the future time range, including information such as steering angle and acceleration;
[0211] The action distribution is the action probability distribution of the ego vehicle at each time step within the future time range;
[0212] The reward sequence is the immediate reward of the ego vehicle at each time step, measuring the safety, comfort, and efficiency of the actions;
[0213] The policy value sequence is the expected cumulative reward of the ego vehicle at each time step.
[0214] 3. Calculate the advantage value sequence and the return sequence using the generalized advantage estimation method.
[0215] 4. Input the label modality, the predicted future trajectories of all environmental agents, the state sequence of the ego vehicle, and the action sequence of the ego vehicle into the trajectory generator to obtain the new action distribution and the new policy value sequence of the ego vehicle.
[0216] 5. Calculate the MSE loss (Mean Squared Error loss) based on the new policy value sequence and the reward sequence to obtain the MSE loss value (i.e., the third loss value).
[0217] Specifically, the MSE loss is defined as follows:
[0218]
[0219] where V t,new and are the new policy value and the reward respectively.
[0220] 6. Calculate the policy loss based on the new action distribution, the original action distribution, and the advantage value sequence to obtain the policy loss value (i.e., the fourth loss value).
[0221] Specifically, the policy loss is defined as follows:
[0222]
[0223] where the ratio r t is given by the following formula: d t,new and d t are the policy distributions (the mean and standard deviation of the Gaussian distribution) induced by π and π old respectively at time step t, the function Prob(a, d) calculates the probability of the given action α under the distribution d, and A t is the advantage value obtained using the generalized advantage estimation method.
[0224] 7. Calculate the entropy loss based on the new action distribution to obtain the entropy loss value (i.e., the fifth loss value).
[0225] Among them, the calculation of the entropy loss is defined as follows:
[0226]
[0227] where represents the entropy of the policy distribution d.
[0228] 8. Take the above three losses as the total loss L generator = L value + L policy + L entropy, and based on the total loss, use the optimizer to update the weights of the trajectory generator, and continuously iterate this process until the preset iteration end condition is met, obtain the trained trajectory generator, and fix the weights of this trajectory generator.
[0229] It should be noted that during the process of model training for the trajectory generation model, the method of using replicas can also be adopted; specifically, this method can randomly initialize the network parameter weights of the trajectory generator and introduce a replica of the trajectory generator, and the weights of this replica are the same as those of the trajectory generator. During the continuous iteration process of the trajectory generator, copy the parameters of the trajectory generator to the replica of the trajectory generator at regular intervals, so as to ensure that in the case of unsatisfactory training results, the weights of the original model are not affected, and thus the retraining can be carried out quickly.
[0230] In one or more embodiments provided in this specification, after the world model, the trajectory selector, and the trajectory generator are all trained, the deployment can be carried out to realize the driving trajectory prediction, so as to accurately predict the future trajectory of the autonomous driving vehicle.
[0231] It should be noted that the model training method and the trajectory determination method in this specification can adopt Markov decision process modeling, and the Markov decision process can be formalized as a seven-tuple where is the state space; is the action space; is the state transition probability; represents the reward function and is bounded; is the initial state distribution; T is the time horizon, and γ is the discount factor of future rewards; the state-action sequence is defined as τ=(s 0 , a 0 , s 1 , a 1 ,..., s T ), where are the state and action at time step t respectively; the goal of reinforcement learning is to maximize the expected return:
[0232]
[0233] One or more embodiments of this specification provide a model training method. During the process of model training, first, a sample data set can be determined, and the trajectory information of the sample object and other sample objects included in the sample data set can be segmented to obtain the historical trajectories and future trajectories of the sample object and other sample objects. Second, based on the map information included in the sample data set, the historical trajectories of the sample object and other sample objects, a sample state can be constructed, and based on the target feasible route and target action modality corresponding to the future trajectory of the sample object, a modality label can be determined. Finally, using a large number of and high-performance sample data such as the sample state, the future trajectories of the sample object and other sample objects, feasible routes, preset action modalities, and modality labels, a world model, a trajectory selection model, and a trajectory generation model are trained to obtain a world model, a trajectory selection model, and a trajectory generation model that can accurately determine future trajectories, avoiding the problem of inaccurate future trajectories, and thus reducing the serious negative impact caused by inaccurate future trajectories during the process of autonomous driving.
[0234] The following combines the attached Figure 4 , taking the application of the model training method provided in this specification in autonomous driving as an example, to further illustrate the model training method. Among them, Figure 4 FIG. shows a processing flowchart of a model training method provided by an embodiment of this specification, which specifically includes the following steps.
[0235] Step 402: Determine training data.
[0236] Specifically, the way this method determines training data is as follows:
[0237] 1. Randomly extract a batch of samples from different scenarios in the expert driving data to form a sample data set. Among them, the sample data set includes multiple sample data, and each sample data consists of map information (such as road network, traffic lights, etc.), the trajectories of all traffic participants (such as the ego vehicle and other environmental intelligent agents: oncoming vehicles), and the feasible lanes of the ego vehicle (such as straight lanes and left-turn lanes).
[0238] 2. Divide the trajectories of all traffic participants into historical trajectories and future trajectories according to a specified time point. For example, if a certain sample is a continuous driving trajectory for 20 seconds, this driving trajectory can be divided at the 8th second as the specified time point, with 0 - 8 seconds as the historical trajectory and 9 - 20 seconds as the future trajectory.
[0239] 3. Construct an initial state based on the map information and the historical trajectories of all traffic participants.
[0240] 4. Use the future trajectories of all traffic participants as the trajectory labels for training the world model, trajectory selector, and trajectory generator.
[0241] 5. Define modalities: The horizontal multi-modalities are the lanes available for the host vehicle (such as the straight lane and the left-turn lane in 1), and the vertical multi-modalities are the average speed descriptions (such as slow deceleration, normal deceleration, constant speed, acceleration, rapid acceleration, etc.).
[0242] 6. Determine the label modalities, that is, the horizontal multi-modalities (such as the straight lane) and the vertical multi-modalities (such as constant speed) corresponding to the future trajectory of the host vehicle.
[0243] Step 404: Train the world model.
[0244] Specifically, the way to train the world model is as follows:
[0245] First, randomly initialize the network parameter weights of the world model, input the initial state into the world model, and obtain the predicted future trajectories of other environmental agents (excluding the host vehicle) output by the world model.
[0246] Second, calculate the L1 regression loss according to the predicted future trajectories of all environmental agents and the corresponding trajectory labels.
[0247] Finally, according to the L1 regression loss value, use the optimizer to update the weights of the world model, continuously iterate this process until the preset iteration end condition is met, obtain the trained world model, and fix the weights of this world model.
[0248] Step 406: Train the trajectory selector.
[0249] Specifically, the way to train the trajectory selector is as follows:
[0250] First, randomly initialize the network parameter weights of the trajectory selector, input the initial state and all modalities c (in the above example, when there are two horizontal multi-modalities and five vertical multi-modalities, all modalities are 10 modalities) into the trajectory selector, and obtain the probabilities of each modality output by the trajectory selector;
[0251] Second, calculate the cross-entropy loss according to the probabilities of each modality and the label modalities;
[0252] Finally, according to the cross-entropy loss value, use the optimizer to update the weights of the trajectory selector, continuously iterate this process until the preset iteration end condition is met, obtain the trained trajectory selector, and fix the weights of this trajectory selector.
[0253] Step 408: Train the trajectory generator.
[0254] Specifically, the way to train the trajectory generator is as follows:
[0255] 1. Randomly initialize the network parameter weights of the trajectory generator, and introduce a copy of the trajectory generator with the same weights as the trajectory generator.
[0256] 2. Input the initial state into the world model (the trained world model) to obtain the predicted future trajectories of other environmental agents output by the world model.
[0257] 3. Input the initial state, label modality, and the predicted future trajectories of other environmental agents into the trajectory generator to obtain the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the ego vehicle output by the trajectory generator.
[0258] Among them, the state sequence is the state of the ego vehicle within the future time range (such as the above-mentioned 12 seconds), including information such as position, speed, and direction;
[0259] The action sequence is the actions of the ego vehicle within the future time range, including information such as steering angle and acceleration;
[0260] The action distribution is the action probability distribution of the ego vehicle at each time step within the future time range;
[0261] The reward sequence is the immediate reward of the ego vehicle at each time step, measuring the safety, comfort, and efficiency of the actions;
[0262] The policy value sequence is the expected cumulative reward of the ego vehicle at each time step.
[0263] 4. Use the generalized advantage estimation method to calculate the advantage value sequence and return sequence.
[0264] 5. Input the label modality, the predicted future trajectories of all environmental agents, the state sequence of the ego vehicle, and the action sequence of the ego vehicle into the trajectory generator to obtain the new action distribution and new policy value sequence of the ego vehicle.
[0265] 6. Calculate the MSE loss according to the new policy value sequence and the return sequence.
[0266] 7. Calculate the policy loss according to the new action distribution, the original action distribution, and the advantage value sequence.
[0267] 8. Calculate the entropy loss according to the new action distribution.
[0268] 9. According to the above three losses, use the optimizer to update the weights of the trajectory generator, and continuously iterate this process until the preset iteration end condition is met, obtain the trained trajectory generator, and fix the weights of this trajectory generator.
[0269] Among them, during the continuous iteration of the trajectory generator, the parameters of the trajectory generator are copied to the copy of the trajectory generator at a certain interval.
[0270] After the model training of the world model, trajectory selector, and trajectory generator using the above steps is all completed, deployment can be carried out to achieve driving trajectory prediction. The specific steps for predicting the future trajectory of an autonomous vehicle using the world model, trajectory selector, and trajectory generator are as follows:
[0271] Suppose the ego vehicle is approaching an intersection, and it is necessary to predict the driving trajectory of the ego vehicle within the next 3 seconds according to the current traffic conditions.
[0272] Step 1: Determine that the trajectory prediction time range is T (e.g., 3 seconds), and the time steps corresponding to this prediction time range are also T (e.g., 3 time steps). For example, in the case where the trajectory prediction time range is T (e.g., 3 seconds), the time steps corresponding to this trajectory prediction time range have a corresponding relationship with the pre-set prediction step size. In the case where the prediction step size is 1 second per step, the time steps corresponding to 3 seconds (i.e., the trajectory prediction time range) are T (i.e., 3 time steps). In the case where the prediction step size is 0.1 second per step, 1 second in the trajectory prediction time range corresponds to 10 time steps. Therefore, the time steps corresponding to the trajectory prediction time range (3 seconds) are 30 time steps.
[0273] Step 2: Iteratively generate the future trajectory of the ego vehicle. First, determine the state at the t-th time step as the current state S t Input the world model and output the future trajectories of other traffic participants (oncoming vehicles and pedestrians) at the (t + 1)-th time step (e.g., the oncoming vehicle will continue to go straight through the intersection within the next 3 seconds, and the pedestrian will stay on the sidewalk within the next 3 seconds).
[0274] where t starts from 0, and the current state S t includes map information (lane lines at the intersection, such as the left-turn lane and straight-through lane at the intersection) and agent information (the current and past trajectories of the ego vehicle and other traffic participants, such as the position and speed of the ego vehicle, the position and speed of oncoming vehicles and pedestrians).
[0275] In addition, a view-invariant module can be used to preprocess the current state S t select the nearest K neighbors, filter feasible lanes, etc., and transform to the ego vehicle coordinate system.
[0276] Step 3: Input the current state S t and all modalities (left-turn lane and low speed, left-turn lane and high speed, straight-through lane and low speed, straight-through lane and high speed) into the trajectory selector, and output the optimal modality (e.g., the probability of left-turn lane and low speed is 0.4, the probability of left-turn lane and high speed is 0.3, the probability of straight-through lane and low speed is 0.2, the probability of straight-through lane and high speed is 0.1; the optimal modality is: left-turn lane and low speed).
[0277] Step 4: Input the current state S t , the preferred modality, and the future trajectories of other traffic participants at the (t + 1)-th time step into the trajectory generator, and output the action a0 of the host vehicle at the t-th time step.
[0278] For example, at time t, the vehicle is located at the coordinate (0, 0), and the action at time t is to move forward 5 meters (which can be imagined as stepping on the accelerator). Then, in this ideal case, the position of the vehicle at time (t + 1) will be predicted to have moved forward 5 meters relative to the position at time t, that is, the new position may be (0, 5).
[0279] Step 5: According to the action a0 of the host vehicle at the t-th time step and the future trajectories of other traffic participants at time step (t + 1),
[0280] wherein, the current and past trajectories (i.e., position and speed) of the host vehicle are updated according to the action a0 at the t-th time step; the current and past trajectories (i.e., position and speed) of other traffic participants are updated according to the future trajectories at time step (t + 1).
[0281] In addition, save the action a 0 of the host vehicle at the t-th time step to the action sequence, and save the states at the t-th time step and the (t + 1)-th time step to the state sequence.
[0282] Step 6: Continue to execute Steps 6.2 - 6.5 until all time steps are iterated.
[0283] Step 7: Obtain the future trajectory of the host vehicle according to the action sequence and state sequence of the host vehicle at all time steps: generate a corresponding driving instruction according to the future trajectory of the host vehicle, so that the host vehicle moves forward based on the driving instruction.
[0284] Based on the above steps, it can be known that the model training method of this specification provides a multi-modal trajectory planning method for autonomous driving reinforcement learning based on a consistent autoregressive model. Through this method, the future trajectory can be accurately predicted during the process of autonomous driving, thereby ensuring the safety and comfort of autonomous driving in various real scenarios.
[0285] Corresponding to the above method embodiment, this specification also provides an embodiment of a trajectory determination device, Figure 5 showing a schematic structural diagram of a trajectory determination device provided by an embodiment of this specification. As Figure 5 shown, the device includes:
[0286] A time determination module 502, configured to determine the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range;
[0287] A state acquisition module 504, configured to acquire the state of the target object at the current time step, where the state includes map information corresponding to the target object, environmental information of the target object, and trajectory information of other objects, and the current time step starts from 0;
[0288] An action acquisition module 506, configured to iteratively acquire the actions and states of the target object at all time steps by using a world model, a trajectory selection model, and a trajectory generation model according to the state at the current time step and a preset action modality;
[0289] A future trajectory determination module 508, configured to determine the future trajectory of the target object according to the actions and states of the target object at all time steps.
[0290] Optionally, the action acquisition module 506 is further configured to:
[0291] Generate an action of the target object at the current time step by using a world model, a trajectory selection model, and a trajectory generation model according to the state at the current time step and a preset action modality;
[0292] In the case of determining that the current time step is not the last time step, determine the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of other objects at the next time step;
[0293] Take the next time step as the current time step, and continue to execute the step of acquiring the state of the target object at the current time step;
[0294] Until, in the case of determining that the current time step is the last time step, acquire the actions and states of the target object at all time steps.
[0295] Optionally, the action acquisition module 506 is further configured to:
[0296] Input the state at the current time step into the world model to obtain the trajectory information of other objects at the next time step;
[0297] Input the state at the current time step and a preset action modality into the trajectory selection model to obtain the target modality of the target object at the next time step;
[0298] Input the current time step, the trajectory information of other objects at the next time step, and the target modality into the trajectory generation model to obtain the action of the target object at the current time step.
[0299] Optionally, the action acquisition module 506 is further configured to:
[0300] Input the state of the target object at the current time step and the target modality into the trajectory generation model to obtain the action of the target object at the current time step.
[0301] When it is determined that the current time step is not the last time step, determine the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of the other objects at the next time step.
[0302] Take the next time step as the current time step and continue to execute the step of obtaining the state of the target object at the current time step.
[0303] Optionally, the future trajectory determination module 508 is configured to:
[0304] Perform sequence combination on the actions and states of the target object at all time steps to obtain the future trajectory of the target object.
[0305] Optionally, the trajectory determination device further includes a data preprocessing module, which is configured to:
[0306] Use the perspective invariant module to perform data preprocessing on the map information corresponding to the target object, the environmental information of the target object, and the trajectory information of other objects, and obtain the preprocessed state of the target object at the current time step.
[0307] One or more embodiments of this specification provide a trajectory determination device. In the process of determining the future trajectory, first, the trajectory prediction time range of the target object and the time steps corresponding to the trajectory prediction time range can be determined, and the state of the target object at the current time step can be obtained; second, three models with different capabilities, namely the world model, the trajectory selection model, and the trajectory generation model, can be used to perform targeted and accurate iterative processing on the state and preset action modality at the current time step, so as to iteratively obtain the actions and states of the target object at all time steps. Finally, based on the actions and states of the target object at all time steps, an accurate future trajectory can be determined for the target object, thereby avoiding the problem of inaccurate future trajectories, and further reducing the serious negative impact caused by inaccurate future trajectories during the process of autonomous driving.
[0308] The above is a schematic solution of a trajectory determination device in this embodiment. It should be noted that the technical solution of this trajectory determination device and the technical solution of the above trajectory determination method belong to the same concept. For the details not described in the technical solution of the trajectory determination device, reference can be made to the description of the technical solution of the above trajectory determination method.
[0309] Corresponding to the above method embodiments, this specification also provides embodiments of a model training apparatus. Figure 6 FIG. Figure 6 shows a schematic structural diagram of a model training apparatus provided by an embodiment of this specification. As Figure 6 shown, the apparatus includes:
[0310] A sample determination module 602, configured to determine a sample data set, where the sample data set includes sample data of multiple different scenarios, and each sample data includes map information corresponding to a sample object, trajectory information of the sample object and a feasible route, and trajectory information of other sample objects;
[0311] A trajectory segmentation module 604, configured to segment the trajectory information of the sample object and the other sample objects to obtain historical trajectories and future trajectories of the sample object and the other sample objects;
[0312] A state construction module 606, configured to construct a sample state according to the map information, historical trajectories of the sample object and the other sample objects;
[0313] A modality label determination module 608, configured to determine a modality label according to a target feasible route and a target action modality corresponding to the future trajectory of the sample object, where the target feasible route is any one of the feasible routes, and the target action modality is any one of preset action modalities;
[0314] A model training module 610, configured to train and obtain a world model, a trajectory selection model, and a trajectory generation model according to the sample state, future trajectories of the sample object and the other sample objects, the feasible route, the preset action modality, and the modality label.
[0315] Optionally, the trajectory segmentation module 604 is configured to:
[0316] Segment the trajectory information of the sample object and the other sample objects according to a preset segmentation time point to obtain historical trajectories and future trajectories of the sample object, historical trajectories and future trajectories of the other sample objects.
[0317] Optionally, the model training module 610 is further configured to:
[0318] Train and obtain the world model according to the sample state;
[0319] Train and obtain the trajectory selection model according to the sample state, the feasible route, the preset action modality, and the modality label;
[0320] Train the trajectory generation model based on the sample state, the modality label, and the world model.
[0321] Optionally, the model training module 610 is further configured to:
[0322] Input the sample state into the initial world model to obtain the predicted future trajectories of the other sample objects;
[0323] Calculate a first loss value based on the predicted future trajectories of the other sample objects and the future trajectories of the other sample objects;
[0324] Update the network parameter weights of the initial world model using an optimizer according to the first loss value until a preset iteration end condition is met, and train to obtain the world model.
[0325] Optionally, the model training module 610 is further configured to:
[0326] Determine an action modality set according to the feasible route and the preset action modality;
[0327] Input the sample state and the action modalities in the action modality set into the initial trajectory selection model to obtain the probabilities of the action modalities in the action modality set;
[0328] Calculate a second loss value according to the probabilities of the action modalities in the action modality set and the modality label;
[0329] Update the network parameter weights of the initial trajectory selection model using an optimizer according to the second loss value until a preset iteration end condition is met, and train to obtain the trajectory selection model.
[0330] Optionally, the model training module 610 is further configured to:
[0331] Input the sample state into the world model to obtain the predicted future trajectories of the other sample objects;
[0332] Input the sample state, the modality label, and the predicted future trajectories of the other sample objects into the initial trajectory generation model to obtain the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the sample object;
[0333] Calculate a third loss value, a fourth loss value, and a fifth loss value according to the state sequence, action sequence, action distribution, reward sequence, and policy value sequence of the sample object;
[0334] Based on the third loss value, the fourth loss value, and the fifth loss value, update the network parameter weights of the initial trajectory generation model using an optimizer until a preset iteration end condition is met, and train to obtain the trajectory generation model.
[0335] Optionally, the model training device further includes a data preprocessing module configured to:
[0336] Use a perspective invariant module to perform data preprocessing on the map information, the sample object, and the historical trajectories of the other sample objects in the sample state to obtain the preprocessed sample state.
[0337] One or more embodiments of this specification provide a model training device. During the model training process, first, a sample data set can be determined, and the trajectory information of the sample object and other sample objects included in the sample data set can be segmented to obtain the historical trajectories and future trajectories of the sample object and other sample objects; second, a sample state can be constructed based on the map information, the historical trajectories of the sample object and other sample objects included in the sample data set, and a modal label can be determined according to the target feasible route and target action modality corresponding to the future trajectory of the sample object; finally, using a large number of and high-performance sample data such as the sample state, the future trajectories of the sample object and other sample objects, the feasible route, the preset action modality, and the modal label, train to obtain a world model, a trajectory selection model, and a trajectory generation model, so as to obtain a world model, a trajectory selection model, and a trajectory generation model that can accurately determine the future trajectory, avoid the problem of inaccurate future trajectories, and further reduce the serious negative impact caused by inaccurate future trajectories during the process of autonomous driving.
[0338] The above is a schematic solution of a model training device according to this embodiment. It should be noted that the technical solution of this model training device and the technical solution of the above model training method belong to the same concept. For the details not described in the technical solution of the model training device, reference can be made to the description of the technical solution of the above model training method.
[0339] Figure 7 FIG. shows a structural block diagram of a computing device 700 according to an embodiment of this specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.
[0340] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0341] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 7 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0342] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0343] Wherein, the processor 720 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above trajectory determination method or model training method are implemented.
[0344] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above trajectory determination method or model training method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solutions of the above trajectory determination method or model training method.
[0345] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the above trajectory determination method or model training method.
[0346] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above trajectory determination method or model training method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solutions of the above trajectory determination method or model training method.
[0347] An embodiment of this specification also provides a computer program product including computer programs / instructions, which, when executed by a processor, implement the steps of the above trajectory determination method or model training method.
[0348] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above trajectory determination method or model training method belong to the same concept. For the detailed content not described in the technical solution of the computer program product, reference can be made to the description of the technical solutions of the above trajectory determination method or model training method.
[0349] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0350] The computer program / instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0351] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0352] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0353] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A trajectory determination method, comprising: Determine a trajectory prediction time range of a target object and a time step corresponding to the trajectory prediction time range; Obtaining the state of the target object at the current time step, wherein the state includes map information corresponding to the target object, environmental information of the target object, and trajectory information of other objects, and the current time step starts from 0; According to the state of the current time step and the preset action mode, the action and state of the target object at all time steps are iteratively obtained by using the world model, the trajectory selection model and the trajectory generation model; The future trajectory of the target object is determined based on the actions and states of the target object at all time steps.
2. The trajectory determination method according to claim 1, wherein the action and state of the target object at all time steps are iteratively obtained by using a world model, a trajectory selection model, and a trajectory generation model according to the state of the current time step and the preset action mode, comprising: Generate the action of the target object at the current time step according to the state of the current time step and the preset action mode by using the world model, the trajectory selection model and the trajectory generation model; In the case where it is determined that the current time step is not the last time step, determining the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of the other objects at the next time step; Taking the next time step as the current time step, and continuing to execute the step of obtaining the state of the target object at the current time step; Until it is determined that the current time step is the last time step, the action and state of the target object in all time steps are obtained.
3. The trajectory determination method according to claim 2, wherein the step of generating the action of the target object at the current time step by using a world model, a trajectory selection model, and a trajectory generation model according to the state of the current time step and the preset action mode comprises: Inputting the state of the current time step into the world model to obtain the trajectory information of the other objects at the next time step; Inputting the state of the current time step and the preset action mode into a trajectory selection model to obtain the target mode of the target object at the next time step; The current time step, the trajectory information of the other objects at the next time step, and the target mode are input into a trajectory generation model to obtain the action of the target object at the current time step.
4. The trajectory determination method according to claim 3, wherein the step of continuing to execute the step of obtaining the state of the target object at the current time step comprises: Inputting the state of the target object at the current time step and the target mode into the trajectory generation model to obtain the action of the target object at the current time step; In the case where it is determined that the current time step is not the last time step, determining the state of the target object at the next time step according to the action of the target object at the current time step and the trajectory information of the other objects at the next time step; The next time step is used as the current time step, and the step of obtaining the state of the target object at the current time step is continued.
5. A model training method, comprising: Determine a sample data set, wherein the sample data set includes a plurality of sample data of different scenarios, each sample data includes map information corresponding to a sample object, trajectory information and a feasible route of the sample object, and trajectory information of other sample objects; Segmenting the trajectory information of the sample object and the other sample objects to obtain historical trajectories and future trajectories of the sample object and the other sample objects; Constructing a sample state according to the map information, the sample object and the historical tracks of the other sample objects; Determine a modality label according to a target feasible route and a target action modality corresponding to the future trajectory of the sample object, wherein the target feasible route is any one of the feasible routes, and the target action modality is any one of the preset action modalities; According to the sample state, the future trajectories of the sample object and the other sample objects, the feasible route, the preset action mode, and the modality label, a world model, a trajectory selection model, and a trajectory generation model are trained.
6. The model training method according to claim 5, wherein the training to obtain a world model, a trajectory selection model, and a trajectory generation model based on the sample state, the future trajectory of the sample object and the other sample objects, the feasible route, and the modality label comprises: According to the sample state, training is performed to obtain the world model; According to the sample state, the feasible route, the preset action mode, and the mode label, training is performed to obtain the trajectory selection model; The trajectory generation model is trained according to the sample state, the modality label, and the world model.
7. The model training method according to claim 6, wherein the training to obtain the world model according to the sample state comprises: Inputting the sample state into an initial world model to obtain predicted future trajectories of the other sample objects; Calculating a first loss value according to the predicted future trajectory of the other sample objects and the future trajectory of the other sample objects; According to the first loss value, the network parameter weights of the initial world model are updated using an optimizer until a preset iteration end condition is met, and the world model is obtained through training.
8. The model training method according to claim 6, wherein the training to obtain the trajectory selection model according to the sample state, the feasible route, the preset action mode, and the mode label comprises: Determining an action mode set according to the feasible route and the preset action mode; Inputting the sample state and the action mode in the action mode set into an initial trajectory selection model to obtain the probability of the action mode in the action mode set; Calculating a second loss value according to the probability of the action modality in the action modality set and the modality label; According to the second loss value, the network parameter weights of the initial trajectory selection model are updated using an optimizer until a preset iteration end condition is met, and the trajectory selection model is obtained by training.
9. The model training method according to claim 6, wherein the step of training the trajectory generation model according to the sample state, the modality label, and the world model comprises: Inputting the sample state into the world model to obtain the predicted future trajectory of the other sample objects; Input the sample state, the modal label, and the predicted future trajectory of the other sample objects into an initial trajectory generation model to obtain a state sequence, an action sequence, an action distribution, a reward sequence, and a strategy value sequence of the sample object; Calculate a third loss value, a fourth loss value, and a fifth loss value according to the state sequence, action sequence, action distribution, reward sequence, and strategy value sequence of the sample object; According to the third loss value, the fourth loss value, and the fifth loss value, the network parameter weights of the initial trajectory generation model are updated using an optimizer until a preset iteration end condition is met, and the trajectory generation model is obtained by training.
10. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 9 are implemented.